Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 2 additions & 3 deletions src/content/docs-es/choose-a-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ Si no sabes cuál coger, busca en la primera columna lo que quieres hacer.
| Mover un agente de código en sesiones largas | `glm5.3` | Está pensado para eso. Necesita el tier premium |
| Lo mismo, pero sin el tier premium | `glm5.3-flash` | Mismo contexto de 1M y cuota generosa |
| Que conteste rápido | `qwen3.8-flash` | Menos profundidad, mucha menos espera |
| Pasarle un audio al modelo directamente | `mimo-v2.5` o `mimo-v2.6-flash` | Los dos oyen audio de forma nativa |
| Pasarle un audio al modelo directamente | `mimo-v2.6-flash` | Oye audio de forma nativa |
| Describir o analizar una imagen | `deepseek-v4-flash` | Cualquiera menos `glm5.3` sirve; este es el mejor |
| Probar cosas sin gastar cuota | `gemma4` | No tiene contador de tokens |
| Montar un buscador o un RAG | `qwen3-embedding` y después `rerank` | Primero recuperas por similitud, luego reordenas por relevancia |
Expand All @@ -45,7 +45,6 @@ Si no sabes cuál coger, busca en la primera columna lo que quieres hacer.
| `glm5.3` | Agentes de código y tareas largas | 1M | texto | 3B tokens/periodo de facturación |
| `glm5.3-flash` | Agentes de código, sin premium | 1M | texto · imagen | 2B tokens/mes |
| `qwen3.8-flash` | Respuestas rápidas | 262K | texto · imagen | 500M tokens/mes |
| `mimo-v2.5` | Audio de entrada, omnimodal | 1M | texto · imagen · audio | 1.0B tokens/mes |
| `mimo-v2.6-flash` | El MiMo más nuevo, omnimodal | 1M | texto · imagen · audio | 1.0B tokens/mes |
| `gemma4` | Tareas cortas y pruebas | 262K | texto · imagen | sin contador |
| `qwen3.6` | Generación anterior | 262K | texto · imagen | sin contador |
Expand All @@ -67,7 +66,7 @@ Las fichas completas, con parámetros, licencias y modos de razonamiento, están

- **El id no es el nombre comercial.** El modelo que en su casa se llama "GLM 5.3 Flash" aquí es `glm5.3-flash`, en minúsculas, sin espacios y con el punto de la versión.
- **`-flash` significa rápido**, no pequeño ni peor: son variantes optimizadas para latencia.
- **El punto de la versión cuenta.** `qwen3.6` y `qwen3.8-flash` son modelos distintos, y `mimo-v2.5` lleva el punto donde lo lleva.
- **El punto de la versión cuenta.** `qwen3.6` y `qwen3.8-flash` son modelos distintos, y `mimo-v2.6-flash` lleva el punto donde lo lleva.
- **Los ids no cambian de significado.** Cuando servimos una variante nueva de un modelo mantenemos su id si la API es la misma. `deepseek-v4-flash`, por ejemplo, pasó a leer imágenes sin cambiar de nombre.
- **Los ids viejos no se apagan de golpe.** `qwen3.6` sigue respondiendo para que las configuraciones que ya lo nombran no se rompan, pero no es lo que te conviene si empiezas hoy.

Expand Down
8 changes: 4 additions & 4 deletions src/content/docs-es/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -343,7 +343,7 @@ console.log(result.language); // "en"
console.log(result.duration); // 5.2
```

## model: mimo-v2.5
## model: mimo-v2.6-flash

omnimodal: chat, visión y audio

Expand All @@ -354,7 +354,7 @@ curl https://api.nan.builders/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-key-here" \
-d '{
"model": "mimo-v2.5",
"model": "mimo-v2.6-flash",
"messages": [{"role": "user", "content": "Hello, how are you?"}],
"max_tokens": 500
}'
Expand All @@ -369,7 +369,7 @@ curl https://api.nan.builders/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-key-here" \
-d '{
"model": "mimo-v2.5",
"model": "mimo-v2.6-flash",
"messages": [{
"role": "user",
"content": [
Expand All @@ -392,7 +392,7 @@ client = OpenAI(
)

response = client.chat.completions.create(
model="mimo-v2.5",
model="mimo-v2.6-flash",
messages=[{
"role": "user",
"content": [
Expand Down
29 changes: 1 addition & 28 deletions src/content/docs-es/models.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -115,33 +115,6 @@ OpenAI y la misma `base URL`.
]}
/>

<ModelCard
id="mimo-v2-5"
name="mimo-v2.5"
tag="310B-15B"
leftLabel="omnimodal: texto, visión y audio"
rightLabel="capacidades"
description="Modelo MoE de 310B parámetros (15B activos), omnimodal de forma nativa con codificadores dedicados de visión y audio. Contexto de 1M tokens. Tool calling y razonamiento. Cuota de 1.0B tokens al mes por miembro. Licencia MIT."
specs={[
{ label: 'Tipo', value: 'MoE (310B total · 15B activos)' },
{ label: 'Cuantización', value: 'FP8' },
{ label: 'Contexto', value: '1M tokens' },
{ label: 'Respuesta máxima', value: '131K tokens' },
{ label: 'Modalidades de entrada', value: 'texto · imagen · audio' },
{ label: 'Modalidades de salida', value: 'texto' },
{ label: 'Cuota mensual', value: '1.0B tokens / miembro' },
{ label: 'Licencia', value: 'MIT' },
]}
items={[
'Tool calling (function calling)',
'Modo razonamiento (se recomienda <code>max_tokens &ge; 300</code>)',
'Visión (entrada de imagen)',
'Audio (entrada de audio)',
'Contexto de 1M tokens',
'Generación en streaming (SSE)',
]}
/>

<ModelCard
id="mimo-v2-6-flash"
name="mimo-v2.6-flash"
Expand Down Expand Up @@ -366,7 +339,7 @@ Un valor que un modelo no puede aplicar nunca es un error.
| `qwen3.6` | `none`, `minimal`, `low`, `medium`, `high`, `max` | `none` y `minimal` se saltan por completo la fase de razonamiento. Los otros cuatro la acotan: low 2.048, medium 8.192, high 16.384, max 32.768 tokens. |
| `gemma4` | `none`, `minimal`, `low`, `medium`, `high`, `max` | Igual que `qwen3.6`: apagado, o un presupuesto de razonamiento entre 2.048 y 32.768 tokens. |
| `deepseek-v4-flash` | cualquier valor (sin efecto) | El modelo decide por petición cuánto razona; el parámetro nunca cambia eso. |
| `qwen3.8-flash` · `mimo-v2.5` · `mimo-v2.6-flash` | aceptado, profundidad no ajustable | El parámetro se acepta y nunca se rechaza, pero estos modelos gestionan su propia profundidad de razonamiento. |
| `qwen3.8-flash` · `mimo-v2.6-flash` | aceptado, profundidad no ajustable | El parámetro se acepta y nunca se rechaza, pero estos modelos gestionan su propia profundidad de razonamiento. |

Sin parámetro, cada modelo usa su propio valor por defecto (razonamiento
activo en `qwen3.6` y `gemma4`, con presupuesto de 16.384 tokens). Más
Expand Down
7 changes: 1 addition & 6 deletions src/content/docs-es/opencode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -45,11 +45,6 @@ Escribe esto en `~/.config/opencode/opencode.json` para tenerlo en todos tus pro
"limit": { "context": 262144, "output": 32768 },
"modalities": { "input": ["text", "image"], "output": ["text"] }
},
"mimo-v2.5": {
"name": "Xiaomi MiMo V2.5",
"limit": { "context": 1048576, "output": 32768 },
"modalities": { "input": ["text", "image", "audio"], "output": ["text"] }
},
"mimo-v2.6-flash": {
"name": "Xiaomi MiMo V2.6 Flash",
"limit": { "context": 1048576, "output": 32768 },
Expand Down Expand Up @@ -81,7 +76,7 @@ Escribe esto en `~/.config/opencode/opencode.json` para tenerlo en todos tus pro
}
```

Esta es la configuración de los 8 modelos LLM que sirve NaN: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` y `glm5.3`.
Esta es la configuración de los 7 modelos LLM que sirve NaN: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` y `glm5.3`.

`@ai-sdk/openai-compatible` es el adaptador genérico, el que habla con cualquier API con forma de OpenAI. No uses `@ai-sdk/openai` a secas: ese espera la API de OpenAI de verdad.

Expand Down
12 changes: 2 additions & 10 deletions src/content/docs-es/pi.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -60,14 +60,6 @@ En `~/.pi/agent/models.json`:
"contextWindow": 262144,
"maxTokens": 32768
},
{
"id": "mimo-v2.5",
"name": "Xiaomi MiMo V2.5",
"reasoning": true,
"input": ["text", "image"],
"contextWindow": 1048576,
"maxTokens": 32768
},
{
"id": "mimo-v2.6-flash",
"name": "Xiaomi MiMo V2.6 Flash",
Expand Down Expand Up @@ -106,7 +98,7 @@ En `~/.pi/agent/models.json`:
}
```

Esta es la configuración de los 8 modelos LLM que sirve NaN: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` y `glm5.3`. Es el mismo conjunto que escribe el [CLI de NaN](/es/docs/nan-cli), y el mismo que publica la página de [OpenCode](/es/docs/opencode).
Esta es la configuración de los 7 modelos LLM que sirve NaN: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` y `glm5.3`. Es el mismo conjunto que escribe el [CLI de NaN](/es/docs/nan-cli), y el mismo que publica la página de [OpenCode](/es/docs/opencode).

`api: "openai-completions"` es lo que le dice a Pi con qué formato hablar. `maxTokens` es el techo de una respuesta, no el contexto. Es la misma cifra que el bloque de [OpenCode](/es/docs/opencode) publica como `limit.output`, porque responde a la misma pregunta.

Expand Down Expand Up @@ -144,7 +136,7 @@ Pídele algo corto. Si contesta, está saliendo por el clúster. Puedes cambiar
- **La clave va escrita en el fichero.** `models.json` vive en tu carpeta personal, así que no suele acabar en un repositorio, pero tenlo en cuenta si sincronizas tu configuración entre máquinas.
- **`maxTokens` no es el contexto.** Es el techo de cada respuesta. El razonamiento sale de ese mismo presupuesto, así que si pides respuestas largas y razonadas, súbelo.
- **Los modelos que declares son los que verás.** Pi no le pregunta al clúster qué hay disponible: muestra lo que haya en la lista.
- **Los dos MiMo entran sin su audio.** `mimo-v2.5` y `mimo-v2.6-flash`: el esquema de Pi acepta `"text"` e `"image"` en `input` y nada más, y un tercer valor no tumba solo a ese modelo: Pi rechaza el fichero entero, con todos los demás proveedores que haya dentro. Los modelos siguen oyendo audio por la API, pero no desde Pi.
- **`mimo-v2.6-flash` entra sin su audio.** El esquema de Pi acepta `"text"` e `"image"` en `input` y nada más, y un tercer valor no tumba solo a ese modelo: Pi rechaza el fichero entero, con todos los demás proveedores que haya dentro. El modelo sigue oyendo audio por la API, pero no desde Pi.
- **`glm5.3` necesita una clave del tier premium.** Está en la lista para que quien lo tenga lo vea. Sin ese tier la respuesta es un `401`, explicado en [Elige tu modelo](/es/docs/choose-a-model).

</Details>
9 changes: 0 additions & 9 deletions src/content/docs-es/vscode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -76,15 +76,6 @@ Se abre un fichero `chatLanguageModels.json`. Déjalo así:
"maxInputTokens": 229376,
"maxOutputTokens": 32768
},
{
"id": "mimo-v2.5",
"name": "Xiaomi MiMo V2.5",
"url": "https://api.nan.builders/v1/chat/completions",
"toolCalling": true,
"vision": true,
"maxInputTokens": 1015808,
"maxOutputTokens": 32768
},
{
"id": "mimo-v2.6-flash",
"name": "Xiaomi MiMo V2.6 Flash",
Expand Down
5 changes: 2 additions & 3 deletions src/content/docs/choose-a-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ If you do not know which one to pick, look for what you want to do in the first
| Drive a coding agent through long sessions | `glm5.3` | It is built for that. Needs the premium tier |
| The same, but without the premium tier | `glm5.3-flash` | Same 1M context and a generous quota |
| Get an answer fast | `qwen3.8-flash` | Less depth, much less waiting |
| Hand the model an audio file directly | `mimo-v2.5` or `mimo-v2.6-flash` | Both hear audio natively |
| Hand the model an audio file directly | `mimo-v2.6-flash` | It hears audio natively |
| Describe or analyze an image | `deepseek-v4-flash` | Any of them except `glm5.3` will do; this is the best |
| Try things without spending quota | `gemma4` | It has no token counter |
| Build a search engine or a RAG | `qwen3-embedding` and then `rerank` | First you retrieve by similarity, then you reorder by relevance |
Expand All @@ -45,7 +45,6 @@ If you do not know which one to pick, look for what you want to do in the first
| `glm5.3` | Coding agents and long tasks | 1M | text | 3B tokens/billing period |
| `glm5.3-flash` | Coding agents, without premium | 1M | text · image | 2B tokens/month |
| `qwen3.8-flash` | Fast answers | 262K | text · image | 500M tokens/month |
| `mimo-v2.5` | Audio input, omnimodal | 1M | text · image · audio | 1.0B tokens/month |
| `mimo-v2.6-flash` | The newest MiMo, omnimodal | 1M | text · image · audio | 1.0B tokens/month |
| `gemma4` | Short tasks and testing | 262K | text · image | no counter |
| `qwen3.6` | Previous generation | 262K | text · image | no counter |
Expand All @@ -67,7 +66,7 @@ The full spec sheets, with parameters, licenses and reasoning modes, are in [Mod

- **The id is not the commercial name.** The model its makers call "GLM 5.3 Flash" is `glm5.3-flash` here, lowercase, no spaces, and with the version dot.
- **`-flash` means fast**, not small or worse: these are variants optimized for latency.
- **The version dot counts.** `qwen3.6` and `qwen3.8-flash` are different models, and `mimo-v2.5` carries its dot where it carries it.
- **The version dot counts.** `qwen3.6` and `qwen3.8-flash` are different models, and `mimo-v2.6-flash` carries its dot where it carries it.
- **Ids do not change meaning.** When we serve a new variant of a model we keep its id if the API is the same. `deepseek-v4-flash`, for instance, started reading images without changing its name.
- **Old ids are not switched off overnight.** `qwen3.6` still answers so that configurations already naming it do not break, but it is not what you want if you are starting today.

Expand Down
8 changes: 4 additions & 4 deletions src/content/docs/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -343,7 +343,7 @@ console.log(result.language); // "en"
console.log(result.duration); // 5.2
```

## model: mimo-v2.5
## model: mimo-v2.6-flash

omnimodal — chat, vision, and audio

Expand All @@ -354,7 +354,7 @@ curl https://api.nan.builders/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-key-here" \
-d '{
"model": "mimo-v2.5",
"model": "mimo-v2.6-flash",
"messages": [{"role": "user", "content": "Hello, how are you?"}],
"max_tokens": 500
}'
Expand All @@ -369,7 +369,7 @@ curl https://api.nan.builders/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-key-here" \
-d '{
"model": "mimo-v2.5",
"model": "mimo-v2.6-flash",
"messages": [{
"role": "user",
"content": [
Expand All @@ -392,7 +392,7 @@ client = OpenAI(
)

response = client.chat.completions.create(
model="mimo-v2.5",
model="mimo-v2.6-flash",
messages=[{
"role": "user",
"content": [
Expand Down
29 changes: 1 addition & 28 deletions src/content/docs/models.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -115,33 +115,6 @@ with the same `base URL`.
]}
/>

<ModelCard
id="mimo-v2-5"
name="mimo-v2.5"
tag="310B-15B"
leftLabel="omnimodal — text, vision & audio"
rightLabel="capabilities"
description="310B parameter MoE model (15B active), natively omnimodal with dedicated vision and audio encoders. 1M token context. Tool calling and reasoning. 1.0B token monthly quota per member. MIT license."
specs={[
{ label: 'Type', value: 'MoE (310B total · 15B active)' },
{ label: 'Quantization', value: 'FP8' },
{ label: 'Context', value: '1M tokens' },
{ label: 'Max answer', value: '131K tokens' },
{ label: 'Input modalities', value: 'text · image · audio' },
{ label: 'Output modalities', value: 'text' },
{ label: 'Monthly quota', value: '1.0B tokens / member' },
{ label: 'License', value: 'MIT' },
]}
items={[
'Tool calling (function calling)',
'Reasoning mode (recommended <code>max_tokens &ge; 300</code>)',
'Vision (image input)',
'Audio (audio input)',
'1M token context',
'Streaming generation (SSE)',
]}
/>

<ModelCard
id="mimo-v2-6-flash"
name="mimo-v2.6-flash"
Expand Down Expand Up @@ -365,7 +338,7 @@ cannot apply is never an error.
| `qwen3.6` | `none`, `minimal`, `low`, `medium`, `high`, `max` | `none` and `minimal` skip the reasoning phase entirely. The other four cap it: low 2,048, medium 8,192, high 16,384, max 32,768 tokens. |
| `gemma4` | `none`, `minimal`, `low`, `medium`, `high`, `max` | Same as `qwen3.6`: off, or a reasoning budget between 2,048 and 32,768 tokens. |
| `deepseek-v4-flash` | any value (no effect) | The model decides per request how much to reason; the parameter never changes that. |
| `qwen3.8-flash` · `mimo-v2.5` · `mimo-v2.6-flash` | accepted, depth not adjustable | The parameter is accepted and never rejected, but these models manage their own reasoning depth. |
| `qwen3.8-flash` · `mimo-v2.6-flash` | accepted, depth not adjustable | The parameter is accepted and never rejected, but these models manage their own reasoning depth. |

With no parameter, every model uses its own default (reasoning on for
`qwen3.6` and `gemma4`, with a 16,384-token budget). More reasoning costs
Expand Down
7 changes: 1 addition & 6 deletions src/content/docs/opencode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -45,11 +45,6 @@ Write this into `~/.config/opencode/opencode.json` to have it in every project,
"limit": { "context": 262144, "output": 32768 },
"modalities": { "input": ["text", "image"], "output": ["text"] }
},
"mimo-v2.5": {
"name": "Xiaomi MiMo V2.5",
"limit": { "context": 1048576, "output": 32768 },
"modalities": { "input": ["text", "image", "audio"], "output": ["text"] }
},
"mimo-v2.6-flash": {
"name": "Xiaomi MiMo V2.6 Flash",
"limit": { "context": 1048576, "output": 32768 },
Expand Down Expand Up @@ -81,7 +76,7 @@ Write this into `~/.config/opencode/opencode.json` to have it in every project,
}
```

This is the config for the 8 LLM models NaN serves: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` and `glm5.3`.
This is the config for the 7 LLM models NaN serves: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` and `glm5.3`.

`@ai-sdk/openai-compatible` is the generic adapter, the one that speaks to any API shaped like OpenAI's. Do not use plain `@ai-sdk/openai`: that one expects the real OpenAI API.

Expand Down
Loading
Loading