Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion src/content/docs-es/choose-a-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ Si no sabes cuál coger, busca en la primera columna lo que quieres hacer.
| `deepseek-v4-flash` | Chat y razonamiento general | 1M | texto · imagen | 3B tokens/mes |
| `glm5.3` | Agentes de código y tareas largas | 1M | texto | 3B tokens/periodo de facturación |
| `glm5.3-flash` | Agentes de código, sin premium | 1M | texto · imagen | 2B tokens/mes |
| `qwen3.8-flash` | Respuestas rápidas | 262K | texto · imagen | 500M tokens/mes |
| `qwen3.8-flash` | Respuestas rápidas | 1M | texto · imagen | 500M tokens/mes |
| `mimo-v2.6-flash` | El MiMo más nuevo, omnimodal | 1M | texto · imagen · audio | 1.0B tokens/mes |
| `gemma4` | Tareas cortas y pruebas | 262K | texto · imagen | sin contador |
| `qwen3.6` | Generación anterior | 262K | texto · imagen | sin contador |
Expand Down
2 changes: 1 addition & 1 deletion src/content/docs-es/cline.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ Cline lo pregunta porque con un proveedor genérico no tiene forma de averiguarl
|---|---|---|---|
| `glm5.3-flash` | 1000000 | sí | sí |
| `deepseek-v4-flash` | 1000000 | sí | sí |
| `qwen3.8-flash` | 262144 | sí | sí |
| `qwen3.8-flash` | 1000000 | sí | sí |
| `glm5.3` | 1000000 | no | sí |

Si marcas imágenes en un modelo que no las acepta, Cline intentará mandarle capturas y la petición fallará.
Expand Down
4 changes: 2 additions & 2 deletions src/content/docs-es/codex.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ wire_api = "chat"

Tres detalles que importan:

- **`wire_api = "chat"`** hace que Codex use `/chat/completions`. Es lo que quieres: el endpoint `/responses` del clúster contesta de una sola vez en lugar de ir emitiendo la respuesta, así que con `"responses"` verías la respuesta aparecer de golpe al final.
- **`wire_api = "chat"`** hace que Codex use `/chat/completions`. Es lo que quieres con la mayoría de modelos: el endpoint `/responses` del clúster solo va emitiendo la respuesta por partes con `deepseek-v4-flash`, y con los demás contesta de una sola vez, así que con `"responses"` verías la respuesta aparecer de golpe al final.
- **El identificador del proveedor no puede ser `openai`, `ollama` ni `lmstudio`**, que están reservados. Por eso se llama `nan`.
- **`base_url` termina en `/v1`** y nada más. No añadas la ruta del endpoint.

Expand Down Expand Up @@ -74,7 +74,7 @@ O dejar varios proveedores declarados y elegir con `--profile` si prefieres perf

- **Las funciones en la nube de Codex no aplican.** Al declarar un proveedor propio, todo va contra NaN desde tu máquina.
- **El razonamiento se ve distinto según el modelo.** Los modelos del clúster emiten su traza de razonamiento a su manera, y Codex no siempre la presenta como con los modelos de OpenAI.
- **Si cambias `wire_api` a `"responses"`**, la respuesta deja de aparecer poco a poco. No es un cuelgue: es que ese endpoint todavía no emite la respuesta por partes.
- **Si cambias `wire_api` a `"responses"`**, solo `deepseek-v4-flash` sigue mostrando la respuesta poco a poco. Con los demás modelos aparece de golpe al final. No es un cuelgue: es que con esos modelos ese endpoint todavía no emite la respuesta por partes.

</Details>

74 changes: 74 additions & 0 deletions src/content/docs-es/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,6 +76,80 @@ for await (const chunk of stream) {

Instalación: `npm install openai`

### salida estructurada en deepseek-v4-flash

`deepseek-v4-flash` rechaza `response_format` `json_schema` con un `400`, y `json_object` solo garantiza JSON válido, no su forma. Para obtener una salida que siga un esquema, define una única herramienta de tipo función cuyo `parameters` sea tu JSON Schema (raíz `"type": "object"`, `"strict": true`) y fuérzala con `tool_choice`. El modelo devuelve argumentos que siguen el esquema en `choices[0].message.tool_calls[0].function.arguments`, como una cadena JSON. `strict` pide que se ajusten exactamente; se aplica donde el modelo admite decodificación estricta, así que valida los argumentos si tu código depende de ellos.

```bash
curl https://api.nan.builders/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-key-here" \
-d '{
"model": "deepseek-v4-flash",
"messages": [{"role": "user", "content": "Ana García is 34 and lives in Valencia."}],
"tools": [{
"type": "function",
"function": {
"name": "save_person",
"description": "Save the person mentioned in the text.",
"strict": true,
"parameters": {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
"city": {"type": "string"}
},
"required": ["name", "age", "city"],
"additionalProperties": false
}
}
}],
"tool_choice": {"type": "function", "function": {"name": "save_person"}}
}'
# → choices[0].message.tool_calls[0].function.arguments:
# {"name": "Ana García", "age": 34, "city": "Valencia"}
```

```python
import json
from openai import OpenAI

client = OpenAI(
api_key="sk-your-key-here",
base_url="https://api.nan.builders/v1"
)

schema = {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
"city": {"type": "string"}
},
"required": ["name", "age", "city"],
"additionalProperties": False
}

response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Ana García is 34 and lives in Valencia."}],
tools=[{
"type": "function",
"function": {
"name": "save_person",
"description": "Save the person mentioned in the text.",
"strict": True,
"parameters": schema
}
}],
tool_choice={"type": "function", "function": {"name": "save_person"}}
)

person = json.loads(response.choices[0].message.tool_calls[0].function.arguments)
print(person) # {'name': 'Ana García', 'age': 34, 'city': 'Valencia'}
```

## model: qwen3-embedding

embeddings vectoriales
Expand Down
8 changes: 4 additions & 4 deletions src/content/docs-es/models.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ OpenAI y la misma `base URL`.
tag="305B MoE"
leftLabel="generación de texto, chat y visión"
rightLabel="capacidades"
description="Modelo MoE de 305B parámetros, servido en su variante Vision-Exp: acepta imágenes de entrada. Contexto de 1M tokens. Tool calling y razonamiento. Cuota de 3B tokens al mes por miembro. Salida estructurada: response_format json_object es compatible (el prompt debe contener la palabra JSON; si no, se rechaza con un 400), json_schema no y se rechaza con un 400. Para salida restringida a un esquema, usa qwen3.6 o gemma4."
description="Modelo MoE de 305B parámetros, servido en su variante Vision-Exp: acepta imágenes de entrada. Contexto de 1M tokens. Tool calling y razonamiento. Cuota de 3B tokens al mes por miembro. Salida estructurada: response_format json_object es compatible (el prompt debe contener la palabra JSON; si no, se rechaza con un 400), json_schema no y se rechaza con un 400. json_object solo garantiza JSON sintácticamente válido, no su forma: describe el esquema en el prompt. Para salida restringida a un esquema, usa qwen3.6 o gemma4, donde json_schema con strict restringe campos, tipos y claves obligatorias; o, en este modelo, fuerza una única herramienta de tipo función cuyo parameters sea tu esquema (ver Ejemplos). Un max_tokens por debajo de 16384 se sube a 16384 para que quepa el razonamiento, así que un prompt a menos de 16384 tokens de la ventana de 1.048.576 se rechaza con un 400 aunque pidas un max_tokens pequeño (normalmente Context length exceeded, a veces un Invalid request genérico)."
specs={[
{ label: 'Tipo', value: 'MoE (305B total)' },
{ label: 'Cuantización', value: 'FP8' },
Expand Down Expand Up @@ -97,10 +97,10 @@ OpenAI y la misma `base URL`.
tag="125B-6B"
leftLabel="generación de texto, chat y visión"
rightLabel="capacidades"
description="Modelo MoE de 125B parámetros (6B activos), multimodal con visión, tool calling y razonamiento activado por defecto. Contexto de 262K tokens, la ventana nativa del modelo. Cuota de 500M tokens al mes por miembro."
description="Modelo MoE de 125B parámetros (6B activos), multimodal con visión, tool calling y razonamiento activado por defecto. Contexto de 1M tokens (1.048.576 tokens). Cuota de 500M tokens al mes por miembro."
specs={[
{ label: 'Tipo', value: 'MoE (125B total · 6B activos)' },
{ label: 'Contexto', value: '262K tokens' },
{ label: 'Contexto', value: '1M tokens' },
{ label: 'Respuesta máxima', value: '131K tokens' },
{ label: 'Modalidades de entrada', value: 'texto · imagen' },
{ label: 'Modalidades de salida', value: 'texto' },
Expand All @@ -111,7 +111,7 @@ OpenAI y la misma `base URL`.
'Tool calling (formato XML)',
'Modo razonamiento (activado por defecto)',
'Visión (entrada de imagen)',
'Contexto de 262K tokens',
'Contexto de 1M tokens',
'Generación en streaming (SSE)',
]}
/>
Expand Down
8 changes: 4 additions & 4 deletions src/content/docs-es/opencode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ Escribe esto en `~/.config/opencode/opencode.json` para tenerlo en todos tus pro
"models": {
"deepseek-v4-flash": {
"name": "DeepSeek V4 Flash",
"limit": { "context": 1048575, "output": 32768 },
"limit": { "context": 1048576, "output": 32768 },
"modalities": { "input": ["text", "image"], "output": ["text"] }
},
"glm5.3-flash": {
Expand All @@ -42,7 +42,7 @@ Escribe esto en `~/.config/opencode/opencode.json` para tenerlo en todos tus pro
},
"qwen3.8-flash": {
"name": "Qwen 3.8 Flash",
"limit": { "context": 262144, "output": 32768 },
"limit": { "context": 1048576, "output": 32768 },
"modalities": { "input": ["text", "image"], "output": ["text"] }
},
"mimo-v2.6-flash": {
Expand Down Expand Up @@ -108,12 +108,12 @@ Después pega tu API key y pulsa Enter.
## Los límites de contexto

```json
"limit": { "context": 1048575, "output": 32768 }
"limit": { "context": 1048576, "output": 32768 }
```

`limit.context` y `limit.output` son los campos que OpenCode lee. Una versión anterior de esta documentación publicaba `contextWindow`, que no existe en [el esquema de OpenCode](https://opencode.ai/config.json): una clave desconocida no da ningún error que nadie vea, OpenCode se queda con su propia suposición sobre la ventana, y el síntoma es una sesión que compacta demasiado pronto en los modelos de contexto largo.

`limit.context` es la ventana que acepta el proxy, que no siempre es aquella con la que se entrenó el modelo: `qwen3.8-flash` se sirve en sus 262K nativos, no en el 1M extendido con YaRN. `limit.output` es un presupuesto del cliente, no un tope del servidor, así que súbelo si necesitas respuestas más largas.
`limit.context` es la ventana que acepta el proxy: `qwen3.8-flash` se sirve con la ventana completa de 1.048.576 tokens, igual que los demás modelos de 1M. `limit.output` es un presupuesto del cliente, no un tope del servidor, así que súbelo si necesitas respuestas más largas.

## El bloque de compactación

Expand Down
4 changes: 2 additions & 2 deletions src/content/docs-es/pi.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ En `~/.pi/agent/models.json`:
"name": "DeepSeek V4 Flash",
"reasoning": true,
"input": ["text", "image"],
"contextWindow": 1048575,
"contextWindow": 1048576,
"maxTokens": 32768
},
{
Expand All @@ -57,7 +57,7 @@ En `~/.pi/agent/models.json`:
"name": "Qwen 3.8 Flash",
"reasoning": true,
"input": ["text", "image"],
"contextWindow": 262144,
"contextWindow": 1048576,
"maxTokens": 32768
},
{
Expand Down
4 changes: 2 additions & 2 deletions src/content/docs-es/vscode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ Se abre un fichero `chatLanguageModels.json`. Déjalo así:
"url": "https://api.nan.builders/v1/chat/completions",
"toolCalling": true,
"vision": true,
"maxInputTokens": 1015807,
"maxInputTokens": 1015808,
"maxOutputTokens": 32768
},
{
Expand All @@ -73,7 +73,7 @@ Se abre un fichero `chatLanguageModels.json`. Déjalo así:
"url": "https://api.nan.builders/v1/chat/completions",
"toolCalling": true,
"vision": true,
"maxInputTokens": 229376,
"maxInputTokens": 1015808,
"maxOutputTokens": 32768
},
{
Expand Down
2 changes: 1 addition & 1 deletion src/content/docs/choose-a-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ If you do not know which one to pick, look for what you want to do in the first
| `deepseek-v4-flash` | General chat and reasoning | 1M | text · image | 3B tokens/month |
| `glm5.3` | Coding agents and long tasks | 1M | text | 3B tokens/billing period |
| `glm5.3-flash` | Coding agents, without premium | 1M | text · image | 2B tokens/month |
| `qwen3.8-flash` | Fast answers | 262K | text · image | 500M tokens/month |
| `qwen3.8-flash` | Fast answers | 1M | text · image | 500M tokens/month |
| `mimo-v2.6-flash` | The newest MiMo, omnimodal | 1M | text · image · audio | 1.0B tokens/month |
| `gemma4` | Short tasks and testing | 262K | text · image | no counter |
| `qwen3.6` | Previous generation | 262K | text · image | no counter |
Expand Down
2 changes: 1 addition & 1 deletion src/content/docs/cline.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ Cline asks because with a generic provider it has no way of finding out:
|---|---|---|---|
| `glm5.3-flash` | 1000000 | yes | yes |
| `deepseek-v4-flash` | 1000000 | yes | yes |
| `qwen3.8-flash` | 262144 | yes | yes |
| `qwen3.8-flash` | 1000000 | yes | yes |
| `glm5.3` | 1000000 | no | yes |

If you tick images on a model that does not accept them, Cline will try to send it screenshots and the request will fail.
Expand Down
4 changes: 2 additions & 2 deletions src/content/docs/codex.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ wire_api = "chat"

Three details that matter:

- **`wire_api = "chat"`** makes Codex use `/chat/completions`. That is what you want: the cluster's `/responses` endpoint answers in one go instead of streaming, so with `"responses"` you would see the answer appear all at once at the end.
- **`wire_api = "chat"`** makes Codex use `/chat/completions`. That is what you want with most models: the cluster's `/responses` endpoint streams incrementally only on `deepseek-v4-flash`, and on the others it answers in one go, so with `"responses"` you would see the answer appear all at once at the end.
- **The provider identifier cannot be `openai`, `ollama` or `lmstudio`**, which are reserved. That is why it is called `nan`.
- **`base_url` ends at `/v1`** and nothing more. Do not add the endpoint path.

Expand Down Expand Up @@ -74,6 +74,6 @@ Or leave several providers declared and pick with `--profile` if you prefer sepa

- **Codex's cloud features do not apply.** Once you declare a provider of your own, everything goes to NaN from your machine.
- **Reasoning looks different depending on the model.** The cluster's models emit their reasoning trace their own way, and Codex does not always present it the way it does with OpenAI's models.
- **If you change `wire_api` to `"responses"`**, the answer stops appearing gradually. It is not a hang: that endpoint does not stream yet.
- **If you change `wire_api` to `"responses"`**, only `deepseek-v4-flash` keeps showing the answer gradually. With the other models it appears all at once at the end. It is not a hang: on those models that endpoint does not stream yet.

</Details>
74 changes: 74 additions & 0 deletions src/content/docs/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,6 +76,80 @@ for await (const chunk of stream) {

Install: `npm install openai`

### structured output on deepseek-v4-flash

`deepseek-v4-flash` rejects `response_format` `json_schema` with a `400`, and `json_object` only guarantees valid JSON, not its shape. To get output that follows a schema, define one function tool whose `parameters` is your JSON Schema (root `"type": "object"`, `"strict": true`) and force it with `tool_choice`. The model returns arguments that follow the schema in `choices[0].message.tool_calls[0].function.arguments`, as a JSON string. `strict` asks for an exact match; it is enforced where the model supports strict decoding, so validate the arguments if your code depends on them.

```bash
curl https://api.nan.builders/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-key-here" \
-d '{
"model": "deepseek-v4-flash",
"messages": [{"role": "user", "content": "Ana García is 34 and lives in Valencia."}],
"tools": [{
"type": "function",
"function": {
"name": "save_person",
"description": "Save the person mentioned in the text.",
"strict": true,
"parameters": {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
"city": {"type": "string"}
},
"required": ["name", "age", "city"],
"additionalProperties": false
}
}
}],
"tool_choice": {"type": "function", "function": {"name": "save_person"}}
}'
# → choices[0].message.tool_calls[0].function.arguments:
# {"name": "Ana García", "age": 34, "city": "Valencia"}
```

```python
import json
from openai import OpenAI

client = OpenAI(
api_key="sk-your-key-here",
base_url="https://api.nan.builders/v1"
)

schema = {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
"city": {"type": "string"}
},
"required": ["name", "age", "city"],
"additionalProperties": False
}

response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Ana García is 34 and lives in Valencia."}],
tools=[{
"type": "function",
"function": {
"name": "save_person",
"description": "Save the person mentioned in the text.",
"strict": True,
"parameters": schema
}
}],
tool_choice={"type": "function", "function": {"name": "save_person"}}
)

person = json.loads(response.choices[0].message.tool_calls[0].function.arguments)
print(person) # {'name': 'Ana García', 'age': 34, 'city': 'Valencia'}
```

## model: qwen3-embedding

vector embeddings
Expand Down
Loading
Loading