diff --git a/src/content/docs-es/choose-a-model.md b/src/content/docs-es/choose-a-model.md index b58cb86..50c664f 100644 --- a/src/content/docs-es/choose-a-model.md +++ b/src/content/docs-es/choose-a-model.md @@ -44,7 +44,7 @@ Si no sabes cuál coger, busca en la primera columna lo que quieres hacer. | `deepseek-v4-flash` | Chat y razonamiento general | 1M | texto · imagen | 3B tokens/mes | | `glm5.3` | Agentes de código y tareas largas | 1M | texto | 3B tokens/periodo de facturación | | `glm5.3-flash` | Agentes de código, sin premium | 1M | texto · imagen | 2B tokens/mes | -| `qwen3.8-flash` | Respuestas rápidas | 262K | texto · imagen | 500M tokens/mes | +| `qwen3.8-flash` | Respuestas rápidas | 1M | texto · imagen | 500M tokens/mes | | `mimo-v2.6-flash` | El MiMo más nuevo, omnimodal | 1M | texto · imagen · audio | 1.0B tokens/mes | | `gemma4` | Tareas cortas y pruebas | 262K | texto · imagen | sin contador | | `qwen3.6` | Generación anterior | 262K | texto · imagen | sin contador | diff --git a/src/content/docs-es/cline.mdx b/src/content/docs-es/cline.mdx index 5c62502..9cfdf2b 100644 --- a/src/content/docs-es/cline.mdx +++ b/src/content/docs-es/cline.mdx @@ -45,7 +45,7 @@ Cline lo pregunta porque con un proveedor genérico no tiene forma de averiguarl |---|---|---|---| | `glm5.3-flash` | 1000000 | sí | sí | | `deepseek-v4-flash` | 1000000 | sí | sí | -| `qwen3.8-flash` | 262144 | sí | sí | +| `qwen3.8-flash` | 1000000 | sí | sí | | `glm5.3` | 1000000 | no | sí | Si marcas imágenes en un modelo que no las acepta, Cline intentará mandarle capturas y la petición fallará. diff --git a/src/content/docs-es/codex.mdx b/src/content/docs-es/codex.mdx index f810cd3..3dca3ac 100644 --- a/src/content/docs-es/codex.mdx +++ b/src/content/docs-es/codex.mdx @@ -42,7 +42,7 @@ wire_api = "chat" Tres detalles que importan: -- **`wire_api = "chat"`** hace que Codex use `/chat/completions`. Es lo que quieres: el endpoint `/responses` del clúster contesta de una sola vez en lugar de ir emitiendo la respuesta, así que con `"responses"` verías la respuesta aparecer de golpe al final. +- **`wire_api = "chat"`** hace que Codex use `/chat/completions`. Es lo que quieres con la mayoría de modelos: el endpoint `/responses` del clúster solo va emitiendo la respuesta por partes con `deepseek-v4-flash`, y con los demás contesta de una sola vez, así que con `"responses"` verías la respuesta aparecer de golpe al final. - **El identificador del proveedor no puede ser `openai`, `ollama` ni `lmstudio`**, que están reservados. Por eso se llama `nan`. - **`base_url` termina en `/v1`** y nada más. No añadas la ruta del endpoint. @@ -74,7 +74,7 @@ O dejar varios proveedores declarados y elegir con `--profile` si prefieres perf - **Las funciones en la nube de Codex no aplican.** Al declarar un proveedor propio, todo va contra NaN desde tu máquina. - **El razonamiento se ve distinto según el modelo.** Los modelos del clúster emiten su traza de razonamiento a su manera, y Codex no siempre la presenta como con los modelos de OpenAI. -- **Si cambias `wire_api` a `"responses"`**, la respuesta deja de aparecer poco a poco. No es un cuelgue: es que ese endpoint todavía no emite la respuesta por partes. +- **Si cambias `wire_api` a `"responses"`**, solo `deepseek-v4-flash` sigue mostrando la respuesta poco a poco. Con los demás modelos aparece de golpe al final. No es un cuelgue: es que con esos modelos ese endpoint todavía no emite la respuesta por partes. diff --git a/src/content/docs-es/examples.md b/src/content/docs-es/examples.md index facb411..99edf8e 100644 --- a/src/content/docs-es/examples.md +++ b/src/content/docs-es/examples.md @@ -76,6 +76,80 @@ for await (const chunk of stream) { Instalación: `npm install openai` +### salida estructurada en deepseek-v4-flash + +`deepseek-v4-flash` rechaza `response_format` `json_schema` con un `400`, y `json_object` solo garantiza JSON válido, no su forma. Para obtener una salida que siga un esquema, define una única herramienta de tipo función cuyo `parameters` sea tu JSON Schema (raíz `"type": "object"`, `"strict": true`) y fuérzala con `tool_choice`. El modelo devuelve argumentos que siguen el esquema en `choices[0].message.tool_calls[0].function.arguments`, como una cadena JSON. `strict` pide que se ajusten exactamente; se aplica donde el modelo admite decodificación estricta, así que valida los argumentos si tu código depende de ellos. + +```bash +curl https://api.nan.builders/v1/chat/completions \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer sk-your-key-here" \ + -d '{ + "model": "deepseek-v4-flash", + "messages": [{"role": "user", "content": "Ana García is 34 and lives in Valencia."}], + "tools": [{ + "type": "function", + "function": { + "name": "save_person", + "description": "Save the person mentioned in the text.", + "strict": true, + "parameters": { + "type": "object", + "properties": { + "name": {"type": "string"}, + "age": {"type": "integer"}, + "city": {"type": "string"} + }, + "required": ["name", "age", "city"], + "additionalProperties": false + } + } + }], + "tool_choice": {"type": "function", "function": {"name": "save_person"}} + }' +# → choices[0].message.tool_calls[0].function.arguments: +# {"name": "Ana García", "age": 34, "city": "Valencia"} +``` + +```python +import json +from openai import OpenAI + +client = OpenAI( + api_key="sk-your-key-here", + base_url="https://api.nan.builders/v1" +) + +schema = { + "type": "object", + "properties": { + "name": {"type": "string"}, + "age": {"type": "integer"}, + "city": {"type": "string"} + }, + "required": ["name", "age", "city"], + "additionalProperties": False +} + +response = client.chat.completions.create( + model="deepseek-v4-flash", + messages=[{"role": "user", "content": "Ana García is 34 and lives in Valencia."}], + tools=[{ + "type": "function", + "function": { + "name": "save_person", + "description": "Save the person mentioned in the text.", + "strict": True, + "parameters": schema + } + }], + tool_choice={"type": "function", "function": {"name": "save_person"}} +) + +person = json.loads(response.choices[0].message.tool_calls[0].function.arguments) +print(person) # {'name': 'Ana García', 'age': 34, 'city': 'Valencia'} +``` + ## model: qwen3-embedding embeddings vectoriales diff --git a/src/content/docs-es/models.mdx b/src/content/docs-es/models.mdx index 3e3c770..2f5461a 100644 --- a/src/content/docs-es/models.mdx +++ b/src/content/docs-es/models.mdx @@ -47,7 +47,7 @@ OpenAI y la misma `base URL`. tag="305B MoE" leftLabel="generación de texto, chat y visión" rightLabel="capacidades" - description="Modelo MoE de 305B parámetros, servido en su variante Vision-Exp: acepta imágenes de entrada. Contexto de 1M tokens. Tool calling y razonamiento. Cuota de 3B tokens al mes por miembro. Salida estructurada: response_format json_object es compatible (el prompt debe contener la palabra JSON; si no, se rechaza con un 400), json_schema no y se rechaza con un 400. Para salida restringida a un esquema, usa qwen3.6 o gemma4." + description="Modelo MoE de 305B parámetros, servido en su variante Vision-Exp: acepta imágenes de entrada. Contexto de 1M tokens. Tool calling y razonamiento. Cuota de 3B tokens al mes por miembro. Salida estructurada: response_format json_object es compatible (el prompt debe contener la palabra JSON; si no, se rechaza con un 400), json_schema no y se rechaza con un 400. json_object solo garantiza JSON sintácticamente válido, no su forma: describe el esquema en el prompt. Para salida restringida a un esquema, usa qwen3.6 o gemma4, donde json_schema con strict restringe campos, tipos y claves obligatorias; o, en este modelo, fuerza una única herramienta de tipo función cuyo parameters sea tu esquema (ver Ejemplos). Un max_tokens por debajo de 16384 se sube a 16384 para que quepa el razonamiento, así que un prompt a menos de 16384 tokens de la ventana de 1.048.576 se rechaza con un 400 aunque pidas un max_tokens pequeño (normalmente Context length exceeded, a veces un Invalid request genérico)." specs={[ { label: 'Tipo', value: 'MoE (305B total)' }, { label: 'Cuantización', value: 'FP8' }, @@ -97,10 +97,10 @@ OpenAI y la misma `base URL`. tag="125B-6B" leftLabel="generación de texto, chat y visión" rightLabel="capacidades" - description="Modelo MoE de 125B parámetros (6B activos), multimodal con visión, tool calling y razonamiento activado por defecto. Contexto de 262K tokens, la ventana nativa del modelo. Cuota de 500M tokens al mes por miembro." + description="Modelo MoE de 125B parámetros (6B activos), multimodal con visión, tool calling y razonamiento activado por defecto. Contexto de 1M tokens (1.048.576 tokens). Cuota de 500M tokens al mes por miembro." specs={[ { label: 'Tipo', value: 'MoE (125B total · 6B activos)' }, - { label: 'Contexto', value: '262K tokens' }, + { label: 'Contexto', value: '1M tokens' }, { label: 'Respuesta máxima', value: '131K tokens' }, { label: 'Modalidades de entrada', value: 'texto · imagen' }, { label: 'Modalidades de salida', value: 'texto' }, @@ -111,7 +111,7 @@ OpenAI y la misma `base URL`. 'Tool calling (formato XML)', 'Modo razonamiento (activado por defecto)', 'Visión (entrada de imagen)', - 'Contexto de 262K tokens', + 'Contexto de 1M tokens', 'Generación en streaming (SSE)', ]} /> diff --git a/src/content/docs-es/opencode.mdx b/src/content/docs-es/opencode.mdx index 881fa7a..7df22c7 100644 --- a/src/content/docs-es/opencode.mdx +++ b/src/content/docs-es/opencode.mdx @@ -32,7 +32,7 @@ Escribe esto en `~/.config/opencode/opencode.json` para tenerlo en todos tus pro "models": { "deepseek-v4-flash": { "name": "DeepSeek V4 Flash", - "limit": { "context": 1048575, "output": 32768 }, + "limit": { "context": 1048576, "output": 32768 }, "modalities": { "input": ["text", "image"], "output": ["text"] } }, "glm5.3-flash": { @@ -42,7 +42,7 @@ Escribe esto en `~/.config/opencode/opencode.json` para tenerlo en todos tus pro }, "qwen3.8-flash": { "name": "Qwen 3.8 Flash", - "limit": { "context": 262144, "output": 32768 }, + "limit": { "context": 1048576, "output": 32768 }, "modalities": { "input": ["text", "image"], "output": ["text"] } }, "mimo-v2.6-flash": { @@ -108,12 +108,12 @@ Después pega tu API key y pulsa Enter. ## Los límites de contexto ```json -"limit": { "context": 1048575, "output": 32768 } +"limit": { "context": 1048576, "output": 32768 } ``` `limit.context` y `limit.output` son los campos que OpenCode lee. Una versión anterior de esta documentación publicaba `contextWindow`, que no existe en [el esquema de OpenCode](https://opencode.ai/config.json): una clave desconocida no da ningún error que nadie vea, OpenCode se queda con su propia suposición sobre la ventana, y el síntoma es una sesión que compacta demasiado pronto en los modelos de contexto largo. -`limit.context` es la ventana que acepta el proxy, que no siempre es aquella con la que se entrenó el modelo: `qwen3.8-flash` se sirve en sus 262K nativos, no en el 1M extendido con YaRN. `limit.output` es un presupuesto del cliente, no un tope del servidor, así que súbelo si necesitas respuestas más largas. +`limit.context` es la ventana que acepta el proxy: `qwen3.8-flash` se sirve con la ventana completa de 1.048.576 tokens, igual que los demás modelos de 1M. `limit.output` es un presupuesto del cliente, no un tope del servidor, así que súbelo si necesitas respuestas más largas. ## El bloque de compactación diff --git a/src/content/docs-es/pi.mdx b/src/content/docs-es/pi.mdx index 0fdd824..b90f1c2 100644 --- a/src/content/docs-es/pi.mdx +++ b/src/content/docs-es/pi.mdx @@ -41,7 +41,7 @@ En `~/.pi/agent/models.json`: "name": "DeepSeek V4 Flash", "reasoning": true, "input": ["text", "image"], - "contextWindow": 1048575, + "contextWindow": 1048576, "maxTokens": 32768 }, { @@ -57,7 +57,7 @@ En `~/.pi/agent/models.json`: "name": "Qwen 3.8 Flash", "reasoning": true, "input": ["text", "image"], - "contextWindow": 262144, + "contextWindow": 1048576, "maxTokens": 32768 }, { diff --git a/src/content/docs-es/vscode.mdx b/src/content/docs-es/vscode.mdx index 4bcea22..ff84a13 100644 --- a/src/content/docs-es/vscode.mdx +++ b/src/content/docs-es/vscode.mdx @@ -55,7 +55,7 @@ Se abre un fichero `chatLanguageModels.json`. Déjalo así: "url": "https://api.nan.builders/v1/chat/completions", "toolCalling": true, "vision": true, - "maxInputTokens": 1015807, + "maxInputTokens": 1015808, "maxOutputTokens": 32768 }, { @@ -73,7 +73,7 @@ Se abre un fichero `chatLanguageModels.json`. Déjalo así: "url": "https://api.nan.builders/v1/chat/completions", "toolCalling": true, "vision": true, - "maxInputTokens": 229376, + "maxInputTokens": 1015808, "maxOutputTokens": 32768 }, { diff --git a/src/content/docs/choose-a-model.md b/src/content/docs/choose-a-model.md index 6c02691..ca79455 100644 --- a/src/content/docs/choose-a-model.md +++ b/src/content/docs/choose-a-model.md @@ -44,7 +44,7 @@ If you do not know which one to pick, look for what you want to do in the first | `deepseek-v4-flash` | General chat and reasoning | 1M | text · image | 3B tokens/month | | `glm5.3` | Coding agents and long tasks | 1M | text | 3B tokens/billing period | | `glm5.3-flash` | Coding agents, without premium | 1M | text · image | 2B tokens/month | -| `qwen3.8-flash` | Fast answers | 262K | text · image | 500M tokens/month | +| `qwen3.8-flash` | Fast answers | 1M | text · image | 500M tokens/month | | `mimo-v2.6-flash` | The newest MiMo, omnimodal | 1M | text · image · audio | 1.0B tokens/month | | `gemma4` | Short tasks and testing | 262K | text · image | no counter | | `qwen3.6` | Previous generation | 262K | text · image | no counter | diff --git a/src/content/docs/cline.mdx b/src/content/docs/cline.mdx index 9f22b58..b02e3fd 100644 --- a/src/content/docs/cline.mdx +++ b/src/content/docs/cline.mdx @@ -45,7 +45,7 @@ Cline asks because with a generic provider it has no way of finding out: |---|---|---|---| | `glm5.3-flash` | 1000000 | yes | yes | | `deepseek-v4-flash` | 1000000 | yes | yes | -| `qwen3.8-flash` | 262144 | yes | yes | +| `qwen3.8-flash` | 1000000 | yes | yes | | `glm5.3` | 1000000 | no | yes | If you tick images on a model that does not accept them, Cline will try to send it screenshots and the request will fail. diff --git a/src/content/docs/codex.mdx b/src/content/docs/codex.mdx index afd379b..71149f2 100644 --- a/src/content/docs/codex.mdx +++ b/src/content/docs/codex.mdx @@ -42,7 +42,7 @@ wire_api = "chat" Three details that matter: -- **`wire_api = "chat"`** makes Codex use `/chat/completions`. That is what you want: the cluster's `/responses` endpoint answers in one go instead of streaming, so with `"responses"` you would see the answer appear all at once at the end. +- **`wire_api = "chat"`** makes Codex use `/chat/completions`. That is what you want with most models: the cluster's `/responses` endpoint streams incrementally only on `deepseek-v4-flash`, and on the others it answers in one go, so with `"responses"` you would see the answer appear all at once at the end. - **The provider identifier cannot be `openai`, `ollama` or `lmstudio`**, which are reserved. That is why it is called `nan`. - **`base_url` ends at `/v1`** and nothing more. Do not add the endpoint path. @@ -74,6 +74,6 @@ Or leave several providers declared and pick with `--profile` if you prefer sepa - **Codex's cloud features do not apply.** Once you declare a provider of your own, everything goes to NaN from your machine. - **Reasoning looks different depending on the model.** The cluster's models emit their reasoning trace their own way, and Codex does not always present it the way it does with OpenAI's models. -- **If you change `wire_api` to `"responses"`**, the answer stops appearing gradually. It is not a hang: that endpoint does not stream yet. +- **If you change `wire_api` to `"responses"`**, only `deepseek-v4-flash` keeps showing the answer gradually. With the other models it appears all at once at the end. It is not a hang: on those models that endpoint does not stream yet. diff --git a/src/content/docs/examples.md b/src/content/docs/examples.md index 339a28a..6094bb4 100644 --- a/src/content/docs/examples.md +++ b/src/content/docs/examples.md @@ -76,6 +76,80 @@ for await (const chunk of stream) { Install: `npm install openai` +### structured output on deepseek-v4-flash + +`deepseek-v4-flash` rejects `response_format` `json_schema` with a `400`, and `json_object` only guarantees valid JSON, not its shape. To get output that follows a schema, define one function tool whose `parameters` is your JSON Schema (root `"type": "object"`, `"strict": true`) and force it with `tool_choice`. The model returns arguments that follow the schema in `choices[0].message.tool_calls[0].function.arguments`, as a JSON string. `strict` asks for an exact match; it is enforced where the model supports strict decoding, so validate the arguments if your code depends on them. + +```bash +curl https://api.nan.builders/v1/chat/completions \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer sk-your-key-here" \ + -d '{ + "model": "deepseek-v4-flash", + "messages": [{"role": "user", "content": "Ana García is 34 and lives in Valencia."}], + "tools": [{ + "type": "function", + "function": { + "name": "save_person", + "description": "Save the person mentioned in the text.", + "strict": true, + "parameters": { + "type": "object", + "properties": { + "name": {"type": "string"}, + "age": {"type": "integer"}, + "city": {"type": "string"} + }, + "required": ["name", "age", "city"], + "additionalProperties": false + } + } + }], + "tool_choice": {"type": "function", "function": {"name": "save_person"}} + }' +# → choices[0].message.tool_calls[0].function.arguments: +# {"name": "Ana García", "age": 34, "city": "Valencia"} +``` + +```python +import json +from openai import OpenAI + +client = OpenAI( + api_key="sk-your-key-here", + base_url="https://api.nan.builders/v1" +) + +schema = { + "type": "object", + "properties": { + "name": {"type": "string"}, + "age": {"type": "integer"}, + "city": {"type": "string"} + }, + "required": ["name", "age", "city"], + "additionalProperties": False +} + +response = client.chat.completions.create( + model="deepseek-v4-flash", + messages=[{"role": "user", "content": "Ana García is 34 and lives in Valencia."}], + tools=[{ + "type": "function", + "function": { + "name": "save_person", + "description": "Save the person mentioned in the text.", + "strict": True, + "parameters": schema + } + }], + tool_choice={"type": "function", "function": {"name": "save_person"}} +) + +person = json.loads(response.choices[0].message.tool_calls[0].function.arguments) +print(person) # {'name': 'Ana García', 'age': 34, 'city': 'Valencia'} +``` + ## model: qwen3-embedding vector embeddings diff --git a/src/content/docs/models.mdx b/src/content/docs/models.mdx index 08c6c95..95608dd 100644 --- a/src/content/docs/models.mdx +++ b/src/content/docs/models.mdx @@ -47,7 +47,7 @@ with the same `base URL`. tag="305B MoE" leftLabel="text generation, chat & vision" rightLabel="capabilities" - description="305B parameter MoE model, served as the Vision-Exp variant: it takes images as input. 1M token context. Tool calling and reasoning. 3B token monthly quota per member. Structured output: response_format json_object is supported (the prompt must contain the word JSON, otherwise it is rejected with a 400), json_schema is not and is rejected with a 400. For schema-constrained output, use qwen3.6 or gemma4." + description="305B parameter MoE model, served as the Vision-Exp variant: it takes images as input. 1M token context. Tool calling and reasoning. 3B token monthly quota per member. Structured output: response_format json_object is supported (the prompt must contain the word JSON, otherwise it is rejected with a 400), json_schema is not and is rejected with a 400. json_object guarantees syntactically valid JSON only, not its shape: describe the schema in the prompt. For schema-constrained output, use qwen3.6 or gemma4, where json_schema with strict constrains fields, types and required keys; or, on this model, force one function tool whose parameters is your schema (see Examples). A max_tokens below 16384 is raised to 16384 so the reasoning fits, so a prompt within 16384 tokens of the 1,048,576-token window is rejected with a 400 even with a small max_tokens (usually Context length exceeded, sometimes a generic Invalid request)." specs={[ { label: 'Type', value: 'MoE (305B total)' }, { label: 'Quantization', value: 'FP8' }, @@ -97,10 +97,10 @@ with the same `base URL`. tag="125B-6B" leftLabel="text generation, chat & vision" rightLabel="capabilities" - description="125B parameter MoE model (6B active), multimodal with vision, tool calling and reasoning on by default. 262K token context, the model's native window. 500M token monthly quota per member." + description="125B parameter MoE model (6B active), multimodal with vision, tool calling and reasoning on by default. 1M token context (1,048,576 tokens). 500M token monthly quota per member." specs={[ { label: 'Type', value: 'MoE (125B total · 6B active)' }, - { label: 'Context', value: '262K tokens' }, + { label: 'Context', value: '1M tokens' }, { label: 'Max answer', value: '131K tokens' }, { label: 'Input modalities', value: 'text · image' }, { label: 'Output modalities', value: 'text' }, @@ -111,7 +111,7 @@ with the same `base URL`. 'Tool calling (XML format)', 'Reasoning mode (on by default)', 'Vision (image input)', - '262K token context', + '1M token context', 'Streaming generation (SSE)', ]} /> diff --git a/src/content/docs/opencode.mdx b/src/content/docs/opencode.mdx index 5c78422..db3d627 100644 --- a/src/content/docs/opencode.mdx +++ b/src/content/docs/opencode.mdx @@ -32,7 +32,7 @@ Write this into `~/.config/opencode/opencode.json` to have it in every project, "models": { "deepseek-v4-flash": { "name": "DeepSeek V4 Flash", - "limit": { "context": 1048575, "output": 32768 }, + "limit": { "context": 1048576, "output": 32768 }, "modalities": { "input": ["text", "image"], "output": ["text"] } }, "glm5.3-flash": { @@ -42,7 +42,7 @@ Write this into `~/.config/opencode/opencode.json` to have it in every project, }, "qwen3.8-flash": { "name": "Qwen 3.8 Flash", - "limit": { "context": 262144, "output": 32768 }, + "limit": { "context": 1048576, "output": 32768 }, "modalities": { "input": ["text", "image"], "output": ["text"] } }, "mimo-v2.6-flash": { @@ -108,12 +108,12 @@ Then paste your API key and press Enter. ## The context limits ```json -"limit": { "context": 1048575, "output": 32768 } +"limit": { "context": 1048576, "output": 32768 } ``` `limit.context` and `limit.output` are the fields OpenCode reads. An older version of these docs published `contextWindow`, which is not part of [OpenCode's schema](https://opencode.ai/config.json): an unknown key raises nothing anyone sees, OpenCode simply falls back to its own assumption about the window, and the symptom is a session that compacts far too early on the long-context models. -`limit.context` is the window the proxy accepts, which is not always the window the model was trained with: `qwen3.8-flash` is served at its native 262K, not at the YaRN-extended 1M. `limit.output` is a client-side budget rather than a server cap, so raise it if you need longer answers. +`limit.context` is the window the proxy accepts: `qwen3.8-flash` is served with the full 1,048,576-token window, the same as the other 1M models. `limit.output` is a client-side budget rather than a server cap, so raise it if you need longer answers. ## The compaction block diff --git a/src/content/docs/pi.mdx b/src/content/docs/pi.mdx index a487adf..e63c96d 100644 --- a/src/content/docs/pi.mdx +++ b/src/content/docs/pi.mdx @@ -41,7 +41,7 @@ In `~/.pi/agent/models.json`: "name": "DeepSeek V4 Flash", "reasoning": true, "input": ["text", "image"], - "contextWindow": 1048575, + "contextWindow": 1048576, "maxTokens": 32768 }, { @@ -57,7 +57,7 @@ In `~/.pi/agent/models.json`: "name": "Qwen 3.8 Flash", "reasoning": true, "input": ["text", "image"], - "contextWindow": 262144, + "contextWindow": 1048576, "maxTokens": 32768 }, { diff --git a/src/content/docs/vscode.mdx b/src/content/docs/vscode.mdx index e17833a..292e278 100644 --- a/src/content/docs/vscode.mdx +++ b/src/content/docs/vscode.mdx @@ -55,7 +55,7 @@ A `chatLanguageModels.json` file opens. Leave it like this: "url": "https://api.nan.builders/v1/chat/completions", "toolCalling": true, "vision": true, - "maxInputTokens": 1015807, + "maxInputTokens": 1015808, "maxOutputTokens": 32768 }, { @@ -73,7 +73,7 @@ A `chatLanguageModels.json` file opens. Leave it like this: "url": "https://api.nan.builders/v1/chat/completions", "toolCalling": true, "vision": true, - "maxInputTokens": 229376, + "maxInputTokens": 1015808, "maxOutputTokens": 32768 }, { diff --git a/src/data/modelos.json b/src/data/modelos.json index f7ff973..0b09481 100644 --- a/src/data/modelos.json +++ b/src/data/modelos.json @@ -37,7 +37,7 @@ { "id": "qwen3.8-flash", "by": "Alibaba", - "specs": "125B-6B MoE · 262K context · vision · tool calling · reasoning", + "specs": "125B-6B MoE · 1M context · vision · tool calling · reasoning", "cuota": "500M tokens/mes", "frontier": true }, diff --git a/src/data/openapi.json b/src/data/openapi.json index b378f8f..4c49298 100644 --- a/src/data/openapi.json +++ b/src/data/openapi.json @@ -3,7 +3,7 @@ "info": { "title": "NaN API", "version": "1.0.0", - "description": "Open models on a shared EU inference cluster. Zero logs.\n\nThe NaN API is OpenAI-compatible: predictable, resource-oriented URLs, JSON request and response bodies, and standard HTTP verbs and status codes. Point any OpenAI SDK at our base URL and your existing code keeps working. Change the base URL and the API key, and that's it.\n\nOne schema across every model, so you only learn the API once. Change the `model` field to switch models; everything else stays the same.\n\n- Base URL: `https://api.nan.builders/v1`\n- OpenAPI spec: this document. Import it into Postman, Insomnia, or your own tooling.\n\nIf you use the [Helmcode](https://helmcode.com) enterprise service, the base URL is `https://api.helmcode.com/v1` instead. Every other endpoint is identical.\n\n## Authentication\n\nEvery request authenticates with an API key, sent as a Bearer token:\n\n```\nAuthorization: Bearer $NAN_API_KEY\n```\n\nYou must be a NaN community member. Generate your key from user settings, under \"API Keys\", on the [platform](https://cloud.nan.builders/). The key is personal and non-transferable. Keep it secret: never embed one in client-side code or commit it to source control. Requests must go over HTTPS; calls over plain HTTP fail.\n\n## Making requests\n\nThe API is OpenAI-compatible, so point an official OpenAI SDK at our base URL and change nothing else:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(\n api_key=\"$NAN_API_KEY\",\n base_url=\"https://api.nan.builders/v1\",\n)\n\nresp = client.chat.completions.create(\n model=\"deepseek-v4-flash\",\n messages=[{\"role\": \"user\", \"content\": \"Hello\"}],\n)\nprint(resp.choices[0].message.content)\n```\n\n## Streaming\n\nChat responses can stream token-by-token. Set `\"stream\": true` on `/chat/completions` and the response arrives as Server-Sent Events: each event is a `data:` line carrying a `chat.completion.chunk`, with the new text in `choices[0].delta.content`. A final `data: [DONE]` line ends the stream. Only `/chat/completions` streams incrementally; `/responses` currently emits a single terminal event.\n\n## Rate limits\n\n{{RATE_LIMITS}}\n\nImage endpoints run on their own budget, separate from the model endpoints: 20 requests per minute and 100 requests per month. The usage endpoint is metered separately too: {{USAGE_RATE_LIMIT}} requests per minute per member. Exceed any limit and you get a `429`.\n\n## Errors\n\nNaN uses conventional HTTP status codes: `2xx` on success, `4xx` for a problem with the request (a missing parameter, an invalid key, an unavailable model) and `5xx` for a server-side error. Every error returns a JSON body in the OpenAI shape:\n\n```json\n{\n \"error\": {\n \"message\": \"The model 'foo' does not exist.\",\n \"type\": \"invalid_request_error\",\n \"param\": \"model\",\n \"code\": \"model_not_found\"\n }\n}\n```\n\n`message` is human-readable, `param` names the offending field when applicable, and `code` is a short machine-readable string you can branch on.\n\n| Status | Meaning | `code` |\n| --- | --- | --- |\n| `400` | Invalid or malformed parameter (`param` says which); or content blocked by the safety filter. | `invalid_request_error` · `content_policy_violation` |\n| `401` | Missing or invalid API key, or a key whose tier does not reach the requested model (`glm5.3`): \"This API key does not have access to the requested model\", `type: auth_error`. Measured 2026-09-12. | `invalid_api_key` |\n| `402` | The token allowance is spent on a model that carries one. Not retryable: the counter returns to zero when that model's quota period does, the calendar month for the models counted per month and your billing period for `glm5.3`. | `monthly_cap_reached` |\n| `403` | Your tier can't access this endpoint. Image generation requires inference membership. A model your tier cannot reach answers `401`, not this. | `tier_restricted` |\n| `404` | The requested model doesn't exist. | `model_not_found` |\n| `429` | Rate limit hit (`rpm_limit`, `max_parallel_requests`), the rolling 4h token budget of `glm5.3`, or a quota exhausted. | `rate_limit_exceeded` · `insufficient_quota` · `quota_exceeded` |\n| `500` | Something went wrong on our side (includes upstream model errors). | (none) |\n| `524` | Timeout, typical with large audio files on `/audio/transcriptions`. | (none) |\n\nRetry `429` and `5xx` responses with exponential backoff. Don't retry `400`, `401`, `403`, or `404` blindly: they'll fail the same way every time until you change the request. `402` cannot be fixed by repetition either: it clears when that model's quota period resets.\n\n## Model catalog\n\nEvery endpoint takes a `model` id. Capabilities vary by model:\n\n| Model | Use for | Capabilities |\n| --- | --- | --- |\n| `deepseek-v4-flash` | Chat, vision, reasoning | Streaming, tool calling, reasoning, image input, 1M-token context. 3B tokens/month per member |\n| `mimo-v2.6-flash` | Chat, vision, audio | Streaming, tool calling, reasoning, image input, audio input, 1M-token context. 1.0B tokens/month per member |\n| `qwen3.8-flash` | Chat, vision, agents | Streaming, tool calling, reasoning (on by default), vision, 262K-token context. 500M tokens/month per member |\n| `glm5.3-flash` | Chat, vision, agents | Streaming, tool calling, reasoning, vision, 1M-token context. 2B tokens/month per member |\n| `qwen3.6` | Chat, agents | Streaming, tool calling, vision, reasoning (opt-out, returns `reasoning_content`) |\n| `gemma4` | Chat, vision, agents | Streaming, tool calling, vision, reasoning (opt-in) |\n| `glm5.3` | Coding, long-horizon agents | Streaming, tool calling, reasoning trace, text-only input, 1M-token context. Premium tier only |\n| `qwen3-embedding` | Embeddings | 4096-dimension vectors |\n| `rerank` | RAG reranking | Qwen3-Reranker-8B, 100+ languages |\n| `kokoro` | Text-to-speech | Multiple voices and audio formats |\n| `whisper` | Speech-to-text | Transcription with word/segment timestamps |\n| `flux-2-klein` | Image generation | Text-to-image and image-to-image |\n| `qwen-image-2.1` | Image generation (text→image) | 512-1280 px, 1-4 per request, seed 0-2147483647. 100 images/month per member (shared pool with flux-2-klein) |\n\n`glm5.3` is served only to keys on the GLM 5.3 premium tier; every other model is available to any inference member. Call [List models](#tag/Models) for the exact set available to your key.\n\n## Versioning & compatibility\n\nThe API tracks the OpenAI API surface, so OpenAI SDKs and tools work against `https://api.nan.builders/v1` unchanged. This reference documents the stable public `/v1` endpoints, and we add capabilities without breaking existing fields.", + "description": "Open models on a shared EU inference cluster. Zero logs.\n\nThe NaN API is OpenAI-compatible: predictable, resource-oriented URLs, JSON request and response bodies, and standard HTTP verbs and status codes. Point any OpenAI SDK at our base URL and your existing code keeps working. Change the base URL and the API key, and that's it.\n\nOne schema across every model, so you only learn the API once. Change the `model` field to switch models; everything else stays the same.\n\n- Base URL: `https://api.nan.builders/v1`\n- OpenAPI spec: this document. Import it into Postman, Insomnia, or your own tooling.\n\nIf you use the [Helmcode](https://helmcode.com) enterprise service, the base URL is `https://api.helmcode.com/v1` instead. Every other endpoint is identical.\n\n## Authentication\n\nEvery request authenticates with an API key, sent as a Bearer token:\n\n```\nAuthorization: Bearer $NAN_API_KEY\n```\n\nYou must be a NaN community member. Generate your key from user settings, under \"API Keys\", on the [platform](https://cloud.nan.builders/). The key is personal and non-transferable. Keep it secret: never embed one in client-side code or commit it to source control. Requests must go over HTTPS; calls over plain HTTP fail.\n\n## Making requests\n\nThe API is OpenAI-compatible, so point an official OpenAI SDK at our base URL and change nothing else:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(\n api_key=\"$NAN_API_KEY\",\n base_url=\"https://api.nan.builders/v1\",\n)\n\nresp = client.chat.completions.create(\n model=\"deepseek-v4-flash\",\n messages=[{\"role\": \"user\", \"content\": \"Hello\"}],\n)\nprint(resp.choices[0].message.content)\n```\n\n## Streaming\n\nChat responses can stream token-by-token. Set `\"stream\": true` on `/chat/completions` and the response arrives as Server-Sent Events: each event is a `data:` line carrying a `chat.completion.chunk`, with the new text in `choices[0].delta.content`. A final `data: [DONE]` line ends the stream. `/chat/completions` streams incrementally on every chat model. `/responses` streams incrementally only on `deepseek-v4-flash`; on `qwen3.6` and `gemma4` it holds the whole answer and sends it in one burst at the end.\n\n## Rate limits\n\n{{RATE_LIMITS}}\n\nImage endpoints run on their own budget, separate from the model endpoints: 20 requests per minute and 100 requests per month. The usage endpoint is metered separately too: {{USAGE_RATE_LIMIT}} requests per minute per member. Exceed any limit and you get a `429`.\n\n## Errors\n\nNaN uses conventional HTTP status codes: `2xx` on success, `4xx` for a problem with the request (a missing parameter, an invalid key, an unavailable model) and `5xx` for a server-side error. Every error returns a JSON body in the OpenAI shape:\n\n```json\n{\n \"error\": {\n \"message\": \"The model 'foo' does not exist.\",\n \"type\": \"invalid_request_error\",\n \"param\": \"model\",\n \"code\": \"model_not_found\"\n }\n}\n```\n\n`message` is human-readable, `param` names the offending field when applicable, and `code` is a short machine-readable string you can branch on.\n\n| Status | Meaning | `code` |\n| --- | --- | --- |\n| `400` | Invalid or malformed parameter (`param` says which); content blocked by the safety filter; a `max_tokens` above what the model can generate (a generic \"Invalid request\" message today); or a request that does not fit the model's context window. On the models hosted on the cluster (`qwen3.6`, `gemma4`) the overflow message is \"Context length exceeded for model '...'\" with the model's limit; on the other models it is usually that text, but it can arrive as the generic \"Invalid request. Check your request parameters.\" | `\"400\"` (with `type: invalid_request_error`) · `content_policy_violation` |\n| `401` | Missing or invalid API key, or a key whose tier does not reach the requested model (`glm5.3`): \"This API key does not have access to the requested model\", `type: auth_error`. Measured 2026-09-12. | `invalid_api_key` |\n| `402` | The token allowance is spent on a model that carries one. Not retryable: the counter returns to zero when that model's quota period does, the calendar month for the models counted per month and your billing period for `glm5.3`. | `monthly_cap_reached` |\n| `403` | Your tier can't access this endpoint. Image generation requires inference membership. A model your tier cannot reach answers `401`, not this. | `tier_restricted` |\n| `404` | The requested model doesn't exist. | `model_not_found` |\n| `429` | Rate limit hit (`rpm_limit`, `max_parallel_requests`), the rolling 4h token budget of `glm5.3`, or a quota exhausted. | `rate_limit_exceeded` · `insufficient_quota` · `quota_exceeded` |\n| `500` | Something went wrong on our side (includes upstream model errors). | (none) |\n| `524` | Timeout, typical with large audio files on `/audio/transcriptions`. | (none) |\n\nRetry `429` and `5xx` responses with exponential backoff. Don't retry `400`, `401`, `403`, or `404` blindly (a context overflow is not retried automatically either: shorten the input or pick a model with a larger window): they'll fail the same way every time until you change the request. `402` cannot be fixed by repetition either: it clears when that model's quota period resets.\n\n## Model catalog\n\nEvery endpoint takes a `model` id. Capabilities vary by model:\n\n| Model | Use for | Capabilities |\n| --- | --- | --- |\n| `deepseek-v4-flash` | Chat, vision, reasoning | Streaming, tool calling, reasoning, image input, 1M-token context. 3B tokens/month per member |\n| `mimo-v2.6-flash` | Chat, vision, audio | Streaming, tool calling, reasoning, image input, audio input, 1M-token context. 1.0B tokens/month per member |\n| `qwen3.8-flash` | Chat, vision, agents | Streaming, tool calling, reasoning (on by default), vision, 1M-token context. 500M tokens/month per member |\n| `glm5.3-flash` | Chat, vision, agents | Streaming, tool calling, reasoning, vision, 1M-token context. 2B tokens/month per member |\n| `qwen3.6` | Chat, agents | Streaming, tool calling, vision, reasoning (opt-out, returns `reasoning_content`) |\n| `gemma4` | Chat, vision, agents | Streaming, tool calling, vision, reasoning (opt-in) |\n| `glm5.3` | Coding, long-horizon agents | Streaming, tool calling, reasoning trace, text-only input, 1M-token context. Premium tier only |\n| `qwen3-embedding` | Embeddings | 4096-dimension vectors |\n| `rerank` | RAG reranking | Qwen3-Reranker-8B, 100+ languages |\n| `kokoro` | Text-to-speech | Multiple voices and audio formats |\n| `whisper` | Speech-to-text | Transcription with word/segment timestamps |\n| `flux-2-klein` | Image generation | Text-to-image and image-to-image |\n| `qwen-image-2.1` | Image generation (text→image) | 512-1280 px, 1-4 per request, seed 0-2147483647. 100 images/month per member (shared pool with flux-2-klein) |\n\n`glm5.3` is served only to keys on the GLM 5.3 premium tier; every other model is available to any inference member. Call [List models](#tag/Models) for the exact set available to your key.\n\n## Versioning & compatibility\n\nThe API tracks the OpenAI API surface, so OpenAI SDKs and tools work against `https://api.nan.builders/v1` unchanged. This reference documents the stable public `/v1` endpoints, and we add capabilities without breaking existing fields.", "contact": { "name": "NaN", "url": "https://nan.builders" @@ -176,7 +176,7 @@ "max_tokens": { "type": "integer", "minimum": 1, - "description": "Maximum number of tokens to generate.", + "description": "Maximum number of tokens to generate. A value above what the model can generate is rejected with a generic `400`. A request that does not fit the model's context window is rejected with a `400`: \"Context length exceeded for model '...'\" on `qwen3.6` and `gemma4`, usually the same text (sometimes the generic \"Invalid request\") on the other models. On `deepseek-v4-flash` a smaller value is raised to 16384 so the reasoning fits, so a prompt within 16384 tokens of its 1,048,576-token window is rejected with a `400` even with a small `max_tokens`.", "example": 512 }, "temperature": { @@ -1013,7 +1013,7 @@ "Responses" ], "summary": "Create response", - "description": "Creates a model response using the OpenAI-style Responses API. Models: `qwen3.6`, `gemma4`. Streaming currently emits a single terminal event; for token-by-token streaming, use [Create chat completion](#tag/Chat) with `stream: true`.", + "description": "Creates a model response using the OpenAI-style Responses API. Models: `deepseek-v4-flash`, `qwen3.6`, `gemma4`. On `deepseek-v4-flash` the `output` array carries a `reasoning` item before the `message`.\n\nStreaming (`stream: true`) depends on the model. `deepseek-v4-flash` streams incrementally, with the full event sequence: `response.created`, the reasoning summary deltas, `response.output_text.delta` as the answer is generated, and `response.completed`. On `qwen3.6` and `gemma4` the answer is held until it is complete and then arrives in one burst (`response.output_text.delta` events followed by `response.completed`); for token-by-token streaming on those two, use [Create chat completion](#tag/Chat) with `stream: true`.", "requestBody": { "required": true, "content": { @@ -1028,6 +1028,7 @@ "model": { "type": "string", "enum": [ + "deepseek-v4-flash", "qwen3.6", "gemma4" ], @@ -1049,6 +1050,11 @@ ], "example": "Hello, how are you?" }, + "stream": { + "type": "boolean", + "default": false, + "description": "Stream the response as Server-Sent Events. Incremental only on `deepseek-v4-flash`; see the description above." + }, "max_output_tokens": { "type": "integer", "minimum": 1, @@ -1741,9 +1747,27 @@ "type": "string", "description": "What the function does, used by the model to decide when and how to call it. Be descriptive." }, + "strict": { + "type": "boolean", + "default": false, + "description": "Request that the generated arguments follow `parameters` exactly (fields, types and required keys); enforced where the model supports strict decoding. Combined with a forced `tool_choice`, it is the way to get schema-shaped output on `deepseek-v4-flash`, which returns arguments that follow the schema." + }, "parameters": { "type": "object", - "description": "The parameters the function accepts, described as a JSON Schema object." + "required": [ + "type" + ], + "properties": { + "type": { + "type": "string", + "enum": [ + "object" + ], + "description": "The root type of the schema. Must be `object`." + } + }, + "additionalProperties": true, + "description": "The parameters the function accepts, described as a JSON Schema object. When present, its root must be `\"type\": \"object\"`: a root of any other type, or with no `type`, is rejected with a `400` before the request reaches the model, on every model. A function that takes no arguments can omit `parameters` or send `{\"type\": \"object\", \"properties\": {}}`." } } } @@ -1796,7 +1820,7 @@ } }, "ResponseFormat": { - "description": "Force structured output. `json_object` guarantees syntactically valid JSON; `json_schema` (with `strict: true`) constrains the output to a schema. `json_schema` works on `qwen3.6` and `gemma4`. On `deepseek-v4-flash` it is rejected with a `400` before the request reaches the model; `json_object` still works there, as long as the prompt contains the word JSON (otherwise `400`).", + "description": "Force structured output. What each mode guarantees:\n\n- `json_object`: syntactically valid JSON only, NOT its shape. Describe the schema you want in the prompt. On `deepseek-v4-flash` the prompt must also contain the word JSON, otherwise `400`.\n- `json_schema` with `strict: true`: the output is constrained to the schema (fields, types and required keys).\n\n`json_schema` works on `qwen3.6` and `gemma4`. On `deepseek-v4-flash` it is rejected with a `400` before the request reaches the model; `json_object` still works there. For output that follows a schema on `deepseek-v4-flash`, define one function tool whose `parameters` is your JSON Schema (root `type: object`, `strict: true`), force it with `tool_choice: {\"type\": \"function\", \"function\": {\"name\": \"...\"}}` and read `choices[0].message.tool_calls[0].function.arguments`.", "oneOf": [ { "type": "object", @@ -2501,7 +2525,7 @@ }, "responses": { "BadRequest": { - "description": "Invalid parameter. The body includes `param` with the offending field. Safety filter returns `content_policy_violation`.", + "description": "Invalid parameter. The body includes `param` with the offending field. Safety filter returns `content_policy_violation`. A request that does not fit the model's context window also answers `400`. On `qwen3.6` and `gemma4` the message is \"Context length exceeded for model '...'\" with the model's limit; on the other models it is usually that text, but it can be the generic \"Invalid request. Check your request parameters.\". It is not retried automatically: shorten the input or pick a model with a larger window. A `max_tokens` above what the model can generate also answers a generic `400`.", "content": { "application/json": { "schema": { diff --git a/src/lib/modelCatalog.ts b/src/lib/modelCatalog.ts index ce96aff..5a041ad 100644 --- a/src/lib/modelCatalog.ts +++ b/src/lib/modelCatalog.ts @@ -108,7 +108,7 @@ export const MODELS: ModelSpec[] = [ id: 'qwen3.8-flash', by: 'Alibaba', kind: 'chat', - contextTokens: 262_144, + contextTokens: 1_000_000, inputs: ['text', 'image'], quota: { kind: 'monthly', label: { en: '500M tokens / mo', es: '500M tokens/mes' } }, endpoint: '/chat/completions', diff --git a/src/lib/openapiSpec.test.ts b/src/lib/openapiSpec.test.ts index bf3db4c..1f5d42e 100644 --- a/src/lib/openapiSpec.test.ts +++ b/src/lib/openapiSpec.test.ts @@ -653,8 +653,9 @@ describe('documented 400s: tool names and structured output', () => { it('says json_schema is rejected on deepseek-v4-flash and json_object works there', () => { const description: string = schemas.ResponseFormat.description; expect(description).toContain( - 'On `deepseek-v4-flash` it is rejected with a `400` before the request reaches the model; `json_object` still works there, as long as the prompt contains the word JSON (otherwise `400`).', + 'On `deepseek-v4-flash` it is rejected with a `400` before the request reaches the model; `json_object` still works there.', ); + expect(description).toContain('On `deepseek-v4-flash` the prompt must also contain the word JSON, otherwise `400`.'); const works = /`json_schema` works on (.+?)\.(?:\s|$)/.exec(description)!; expect(works[1]).not.toContain('deepseek-v4-flash'); }); @@ -679,3 +680,251 @@ describe('documented 400s: tool names and structured output', () => { expect(es).toContain('json_schema no soportado (400)'); }); }); + +/** + * /responses ON deepseek-v4-flash, AND ITS STREAMING, MEASURED 2026-09-29 + * with a real request against api.nan.builders: 200 with a `reasoning` item + * and a `message` item, non-streaming; with `stream: true` it emits the full + * event sequence incrementally (`response.created` first, deltas as they are + * generated, `response.completed` last). qwen3.6 and gemma4 answer 200 too, + * but hold the answer and send every delta in one burst at the end, so the + * old "a single terminal event" sentence was no longer true for any model. + */ +describe('documented /responses: models and streaming', () => { + const op = (spec.paths as any)['/responses'].post; + const body = op.requestBody.content['application/json'].schema.properties; + + it('lists deepseek-v4-flash next to qwen3.6 and gemma4, in text and in the enum', () => { + expect(op.description).toContain('Models: `deepseek-v4-flash`, `qwen3.6`, `gemma4`.'); + expect(body.model.enum).toEqual(['deepseek-v4-flash', 'qwen3.6', 'gemma4']); + for (const id of body.model.enum) expect(NAN_MODELS, id).toContain(id); + }); + + it('says which models stream incrementally and which send one burst', () => { + expect(op.description).toMatch(/`deepseek-v4-flash` streams incrementally/); + expect(op.description).toContain('`response.created`'); + expect(op.description).toMatch(/On `qwen3\.6` and `gemma4` the answer is held until it is complete/); + expect(op.description).toContain('[Create chat completion](#tag/Chat)'); + expect(body.stream.type).toBe('boolean'); + expect(body.stream.default).toBe(false); + }); + + it('drops the stale single-terminal-event claim everywhere', () => { + expect(raw).not.toMatch(/single terminal event/); + expect(spec.info.description).toContain( + '`/responses` streams incrementally only on `deepseek-v4-flash`', + ); + }); + + it('the Codex guide agrees, in both locales', () => { + const page = (locale: string) => + readFileSync( + resolve(dirname(fileURLToPath(import.meta.url)), `../content/docs${locale}/codex.mdx`), + 'utf-8', + ); + expect(page('')).toContain('streams incrementally only on `deepseek-v4-flash`'); + expect(page('')).not.toContain('It is not a hang: that endpoint does not stream yet.'); + expect(page('-es')).toContain('solo va emitiendo la respuesta por partes con `deepseek-v4-flash`'); + }); +}); + +/** + * qwen3.8-flash IS SERVED AT 1,048,576 TOKENS, not 262,144. Re-measured + * 2026-09-29: `max_input_tokens` 1048576 on the community proxy's deployment + * and a 1_048_576 window in the rate-limit hook. The old "262K, the model's + * native window" copy sent clients to compact at a quarter of what the API + * accepts. Other models' figures are untouched and are pinned elsewhere. + */ +describe('documented context: qwen3.8-flash is 1M', () => { + const here = dirname(fileURLToPath(import.meta.url)); + const read = (p: string) => readFileSync(resolve(here, p), 'utf-8'); + + it('the API reference catalog says 1M', () => { + const row = spec.info.description.split('\n').find((l: string) => l.startsWith('| `qwen3.8-flash`')); + expect(row).toContain('1M-token context'); + expect(row).not.toMatch(/262/); + }); + + for (const locale of ['', '-es']) { + it(`no page gives qwen3.8-flash a 262K window (${locale || 'en'})`, () => { + const card = (() => { + const page = read(`../content/docs${locale}/models.mdx`); + const start = page.indexOf('id="qwen3-8-flash"'); + expect(start).toBeGreaterThan(-1); + return page.slice(start, page.indexOf('/>', start)); + })(); + expect(card).not.toMatch(/262|native|nativa/); + expect(card).toMatch(/1M tokens/); + expect(card).toMatch(/1[.,]048[.,]576/); + + expect(read(`../content/docs${locale}/choose-a-model.md`)).toMatch(/^\| `qwen3\.8-flash` \|[^|]+\| 1M \|/m); + expect(read(`../content/docs${locale}/cline.mdx`)).toMatch(/^\| `qwen3\.8-flash` \| 1000000 \|/m); + expect(read(`../content/docs${locale}/opencode.mdx`)).toMatch( + /"name": "Qwen 3\.8 Flash",\s*"limit": \{ "context": 1048576, "output": 32768 \}/, + ); + // VS Code adds input and output, so the input is the window minus 32768. + expect(read(`../content/docs${locale}/vscode.mdx`)).toMatch( + /"id": "qwen3\.8-flash",[^}]*"maxInputTokens": 1015808,\s*"maxOutputTokens": 32768/, + ); + expect(read(`../content/docs${locale}/pi.mdx`)).toMatch( + /"id": "qwen3\.8-flash",[^}]*"contextWindow": 1048576,/, + ); + }); + } + + it('the home table says 1M', () => { + const row = JSON.stringify(modelos).match(/"id":"qwen3\.8-flash"[^}]*"specs":"([^"]+)"/); + expect(row, 'no qwen3.8-flash row in modelos.json').not.toBeNull(); + expect(row![1]).toContain('1M context'); + }); +}); + +/** + * TWO MORE 400s FROM 2026-09-29. A tool whose `parameters` root is not + * `"type": "object"` is rejected before routing, on every model. A request + * that overflows the model's window answers 400 and is not retried + * automatically. The exact "Context length exceeded for model '...'" text was + * measured on the hosted models (gemma4); on deepseek-v4-flash an overflow + * measured 2026-09-29 came back as the GENERIC "Invalid request. Check your + * request parameters.", so the docs promise the text only where measured. On + * deepseek-v4-flash a small `max_tokens` is raised to 16384 so the reasoning + * fits, which is why a prompt close to the window overflows with a small one. + */ +describe('documented 400s: tool parameters and context overflow', () => { + const schemas = spec.components.schemas as any; + const params = schemas.Tool.properties.function.properties.parameters; + + it('pins the object root of tool parameters, machine-readably and in prose', () => { + expect(params.type).toBe('object'); + expect(params.required).toEqual(['type']); + expect(params.properties.type.enum).toEqual(['object']); + // Declaring properties.type must not read as "type is the only field". + expect(params.additionalProperties).toBe(true); + expect(params.description).toContain('its root must be `"type": "object"`'); + expect(params.description).toMatch(/rejected with a `400`/); + expect(params.description).toMatch(/every model/); + // `parameters` stays optional: a no-argument tool may omit it. + expect(schemas.Tool.properties.function.required).toEqual(['name']); + }); + + it('documents the context-overflow 400 in the errors table and the shared 400', () => { + const row = spec.info.description.split('\n').find((l: string) => l.startsWith('| `400` |')); + expect(row).toContain("Context length exceeded for model '...'"); + expect(row).toMatch(/model's limit/); + // The exact text is promised only for the hosted models, where it was measured. + expect(row).toMatch(/hosted on the cluster \(`qwen3\.6`, `gemma4`\) the overflow message is "Context length exceeded/); + expect(row).toContain('Invalid request. Check your request parameters.'); + expect(row).toMatch(/`max_tokens` above what the model can generate/); + // The API returns invalid_request_error in `type`; `code` is "400". + expect(row).toContain('`"400"` (with `type: invalid_request_error`)'); + expect(spec.info.description).not.toMatch(/not retried on another deployment/); + expect(spec.info.description).toContain('not retried automatically either: shorten the input or pick a model with a larger window'); + const bad = (spec.components.responses as any).BadRequest.description; + expect(bad).toContain("Context length exceeded for model '...'"); + expect(bad).toMatch(/not retried automatically: shorten the input or pick a model with a larger window/); + expect(bad).toMatch(/On `qwen3\.6` and `gemma4` the message is "Context length exceeded/); + expect(bad).toContain('Invalid request. Check your request parameters.'); + }); + + it('explains the deepseek-v4-flash 16384 floor on max_tokens, in the spec and on both cards', () => { + const maxTokens = (spec.paths as any)['/chat/completions'].post.requestBody.content['application/json'] + .schema.properties.max_tokens.description; + expect(maxTokens).toContain('On `deepseek-v4-flash` a smaller value is raised to 16384'); + expect(maxTokens).toContain('1,048,576-token window'); + // Not promised as the exact text on deepseek-v4-flash: measured generic there. + expect(maxTokens).not.toMatch(/deepseek-v4-flash[^.]*Context length exceeded/); + expect(maxTokens).toMatch(/above what the model can generate is rejected with a generic `400`/); + + const card = (locale: string) => { + const page = readFileSync( + resolve(dirname(fileURLToPath(import.meta.url)), `../content/docs${locale}/models.mdx`), + 'utf-8', + ); + const start = page.indexOf('id="deepseek-v4-flash"'); + return page.slice(start, page.indexOf('/>', start)); + }; + expect(card('')).toContain('A max_tokens below 16384 is raised to 16384 so the reasoning fits'); + expect(card('')).toContain('rejected with a 400 even with a small max_tokens (usually Context length exceeded, sometimes a generic Invalid request)'); + expect(card('-es')).toContain('Un max_tokens por debajo de 16384 se sube a 16384'); + expect(card('-es')).toContain('se rechaza con un 400 aunque pidas un max_tokens pequeño (normalmente Context length exceeded, a veces un Invalid request genérico)'); + }); +}); + +/** + * WHAT EACH STRUCTURED-OUTPUT MODE GUARANTEES, and the way to get a schema on + * deepseek-v4-flash, where `json_schema` is rejected. `json_object` promises + * valid JSON and nothing about its shape; a member who reads "structured + * output" as "my fields" gets surprised in production. The forced-tool recipe + * was verified live 2026-09-29: both snippets on /docs/examples, run as + * written, returned arguments matching the schema on deepseek-v4-flash. + */ +describe('documented structured output: guarantees and the deepseek-v4-flash recipe', () => { + const here = dirname(fileURLToPath(import.meta.url)); + const read = (p: string) => readFileSync(resolve(here, p), 'utf-8'); + const rf = (spec.components.schemas as any).ResponseFormat.description as string; + + it('ResponseFormat states what json_object and json_schema each guarantee', () => { + expect(rf).toContain('`json_object`: syntactically valid JSON only, NOT its shape. Describe the schema you want in the prompt.'); + expect(rf).toContain('On `deepseek-v4-flash` the prompt must also contain the word JSON'); + expect(rf).toContain('`json_schema` with `strict: true`: the output is constrained to the schema (fields, types and required keys).'); + expect(rf).toMatch(/force it with `tool_choice: \{"type": "function", "function": \{"name": "\.\.\."\}\}`/); + expect(rf).toContain('`choices[0].message.tool_calls[0].function.arguments`'); + }); + + it('Tool.function declares `strict`', () => { + const strict = (spec.components.schemas as any).Tool.properties.function.properties.strict; + expect(strict.type).toBe('boolean'); + expect(strict.description).toMatch(/deepseek-v4-flash/); + // Not a guarantee on every model: decode-time enforcement is per model. + expect(strict.description).toContain('Request that the generated arguments follow `parameters` exactly'); + expect(strict.description).toContain('enforced where the model supports strict decoding'); + expect(strict.description).not.toMatch(/^Constrain/); + }); + + it('the deepseek-v4-flash card says the same, in both locales', () => { + const card = (locale: string) => { + const page = read(`../content/docs${locale}/models.mdx`); + const start = page.indexOf('id="deepseek-v4-flash"'); + return page.slice(start, page.indexOf('/>', start)); + }; + expect(card('')).toContain('json_object guarantees syntactically valid JSON only, not its shape: describe the schema in the prompt'); + expect(card('')).toContain('json_schema with strict constrains fields, types and required keys'); + expect(card('')).toContain('force one function tool whose parameters is your schema'); + expect(card('-es')).toContain('json_object solo garantiza JSON sintácticamente válido, no su forma: describe el esquema en el prompt'); + expect(card('-es')).toContain('json_schema con strict restringe campos, tipos y claves obligatorias'); + expect(card('-es')).toContain('fuerza una única herramienta de tipo función cuyo parameters sea tu esquema'); + }); + + for (const [locale, heading] of [ + ['', '### structured output on deepseek-v4-flash'], + ['-es', '### salida estructurada en deepseek-v4-flash'], + ] as const) { + it(`examples.md carries the forced-tool recipe with curl and python (${locale || 'en'})`, () => { + const page = read(`../content/docs${locale}/examples.md`); + const start = page.indexOf(heading); + expect(start, 'no structured-output section').toBeGreaterThan(-1); + const section = page.slice(start, page.indexOf('\n## ', start)); + const curl = /```bash\n([\s\S]*?)```/.exec(section)?.[1] ?? ''; + const py = /```python\n([\s\S]*?)```/.exec(section)?.[1] ?? ''; + for (const code of [curl, py]) { + expect(code).toMatch(/"?model"?[=:] ?"deepseek-v4-flash"/); + expect(code).toMatch(/"type": "object"/); + expect(code).toMatch(/"strict": (true|True)/); + expect(code).toMatch(/tool_choice"?[=:] ?\{"type": "function", "function": \{"name": "save_person"\}\}/); + expect(code).not.toContain('response_format'); + } + expect(curl).toContain('https://api.nan.builders/v1/chat/completions'); + // The curl body must be the JSON it claims to be. + const body = /-d '([\s\S]*?)'\n/.exec(curl)?.[1]; + expect(body, 'no -d body').toBeDefined(); + const parsed = JSON.parse(body!); + expect(parsed.tools).toHaveLength(1); + expect(parsed.tools[0].function.parameters.type).toBe('object'); + expect(parsed.tool_choice.function.name).toBe(parsed.tools[0].function.name); + expect(py).toContain('response.choices[0].message.tool_calls[0].function.arguments'); + expect(section).toMatch(locale ? /devuelve argumentos que siguen el esquema/ : /returns arguments that follow the schema/); + expect(section).toMatch(locale ? /donde el modelo admite decodificación estricta/ : /where the model supports strict decoding/); + expect(section).not.toMatch(/guaranteed|garantizad/); + }); + } +}); diff --git a/src/tests/lib/docsClientConfigs.test.ts b/src/tests/lib/docsClientConfigs.test.ts index ed8fea4..6ffd51b 100644 --- a/src/tests/lib/docsClientConfigs.test.ts +++ b/src/tests/lib/docsClientConfigs.test.ts @@ -63,23 +63,22 @@ const here = dirname(fileURLToPath(import.meta.url)); * and, for qwen3.6, the live `--max-model-len=262144` on the five * `vllm-qwen36-*` deployments in `nan-inference`: * - * deepseek-v4-flash 1048575 declared glm5.3 1048576 declared - * qwen3.8-flash 262144 declared glm5.3-flash 1048576 declared + * deepseek-v4-flash 1048576 declared glm5.3 1048576 declared (deepseek-v4-flash re-measured 2026-09-29) + * qwen3.8-flash 1048576 declared glm5.3-flash 1048576 declared (qwen3.8-flash re-measured 2026-09-29) * qwen3.6 262144 --max-model-len * gemma4 262144 model card mimo-v2.6-flash 1048576 model card * * EVERY DISAGREEMENT WITH `modelRateLimits`, since the previous version of this * list claimed to be complete and was not: - * * qwen3.8-flash -- 262144 here, 1_048_576 there ("YaRN extends the native - * window to 1M, which is what we expose to members") while the deployment - * declares 262144 and the model card calls 262K "the model's native - * window". Tracked as helmcode/nan#53; it also inflates that model's ITPM - * fourfold. + * * qwen3.8-flash -- RESOLVED. It was 262144 here against 1_048_576 there + * (helmcode/nan#53). Re-measured 2026-09-29: the deployment now declares + * `max_input_tokens` 1048576 and the rate-limit hook windows it at + * 1_048_576, so the configs, the model card and choose-a-model publish 1M. * * mimo-v2.6-flash -- 1048576 here against 1_050_000 there. 1,424 tokens * (inherited verbatim from the retired mimo-v2.5 row). - * * deepseek-v4-flash -- 1048575 here against 1_048_576 there. ONE token, - * and the odd number is the real declaration, corroborated at - * `litellm-community/values.yaml:199`. + * * deepseek-v4-flash -- RESOLVED. It was 1048575 here against 1_048_576 + * there, ONE token. Re-measured 2026-09-29: the deployment now declares + * `max_input_tokens` 1048576, so every config publishes 1048576. * The three GLM groups agree since cloud-api `8aa5496` raised them to 1M. * `output` is NOT a server cap for the self-hosted models -- vLLM bounds the * completion by the context window minus the prompt, with no separate limit. @@ -147,8 +146,8 @@ const ALLOWED_PLACEHOLDERS = new Set([ const EXPECTED_MODELS: Record = { 'qwen3.6': { context: 262_144, output: 65_536 }, gemma4: { context: 262_144, output: 65_536 }, - 'deepseek-v4-flash': { context: 1_048_575, output: 32_768 }, - 'qwen3.8-flash': { context: 262_144, output: 32_768 }, + 'deepseek-v4-flash': { context: 1_048_576, output: 32_768 }, + 'qwen3.8-flash': { context: 1_048_576, output: 32_768 }, 'mimo-v2.6-flash': { context: 1_048_576, output: 32_768 }, 'glm5.3-flash': { context: 1_048_576, output: 32_768 }, };