Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion src/content/docs-es/choose-a-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ Si no sabes cuál coger, busca en la primera columna lo que quieres hacer.
| Mover un agente de código en sesiones largas | `glm5.3` | Está pensado para eso. Necesita el tier premium |
| Lo mismo, pero sin el tier premium | `glm5.3-flash` | Mismo contexto de 1M y cuota generosa |
| Que conteste rápido | `qwen3.8-flash` | Menos profundidad, mucha menos espera |
| Pasarle un audio al modelo directamente | `mimo-v2.5` | Es el único que oye |
| Pasarle un audio al modelo directamente | `mimo-v2.5` o `mimo-v2.6-flash` | Los dos oyen audio de forma nativa |
| Describir o analizar una imagen | `deepseek-v4-flash` | Cualquiera menos `glm5.3` sirve; este es el mejor |
| Probar cosas sin gastar cuota | `gemma4` | No tiene contador de tokens |
| Montar un buscador o un RAG | `qwen3-embedding` y después `rerank` | Primero recuperas por similitud, luego reordenas por relevancia |
Expand All @@ -45,6 +45,7 @@ Si no sabes cuál coger, busca en la primera columna lo que quieres hacer.
| `glm5.3-flash` | Agentes de código, sin premium | 1M | texto · imagen | 2B tokens/mes |
| `qwen3.8-flash` | Respuestas rápidas | 262K | texto · imagen | 500M tokens/mes |
| `mimo-v2.5` | Audio de entrada, omnimodal | 1M | texto · imagen · audio | 1.0B tokens/mes |
| `mimo-v2.6-flash` | El MiMo más nuevo, omnimodal | 1M | texto · imagen · audio | 1.0B tokens/mes |
| `gemma4` | Tareas cortas y pruebas | 262K | texto · imagen | sin contador |
| `qwen3.6` | Generación anterior | 262K | texto · imagen | sin contador |
| `qwen3-embedding` | Vectores de 4096 dimensiones | - | texto | sin contador |
Expand Down
25 changes: 24 additions & 1 deletion src/content/docs-es/models.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,29 @@ OpenAI y la misma `base URL`.
]}
/>

<ModelCard
id="mimo-v2-6-flash"
name="mimo-v2.6-flash"
tag="omnimodal"
leftLabel="omnimodal: texto, visión y audio"
rightLabel="capacidades"
description="El Xiaomi MiMo más nuevo, omnimodal de forma nativa con entrada de visión y audio. Contexto de 1M tokens. Tool calling y razonamiento. Mismos límites que mimo-v2.5: cuota de 1.0B tokens al mes por miembro."
specs={[
{ label: 'Contexto', value: '1M tokens' },
{ label: 'Modalidades de entrada', value: 'texto · imagen · audio' },
{ label: 'Modalidades de salida', value: 'texto' },
{ label: 'Cuota mensual', value: '1.0B tokens / miembro' },
]}
items={[
'Tool calling (function calling)',
'Modo razonamiento (se recomienda <code>max_tokens &ge; 300</code>)',
'Visión (entrada de imagen)',
'Audio (entrada de audio)',
'Contexto de 1M tokens',
'Generación en streaming (SSE)',
]}
/>

<ModelCard
id="gemma4"
name="gemma4"
Expand Down Expand Up @@ -323,7 +346,7 @@ Un valor que un modelo no puede aplicar nunca es un error.
| `qwen3.6` | `none`, `minimal`, `low`, `medium`, `high`, `max` | `none` y `minimal` se saltan por completo la fase de razonamiento. Los otros cuatro la acotan: low 2.048, medium 8.192, high 16.384, max 32.768 tokens. |
| `gemma4` | `none`, `minimal`, `low`, `medium`, `high`, `max` | Igual que `qwen3.6`: apagado, o un presupuesto de razonamiento entre 2.048 y 32.768 tokens. |
| `deepseek-v4-flash` | cualquier valor (sin efecto) | El modelo decide por petición cuánto razona; el parámetro nunca cambia eso. |
| `qwen3.8-flash` · `mimo-v2.5` | aceptado, profundidad no ajustable | El parámetro se acepta y nunca se rechaza, pero estos modelos gestionan su propia profundidad de razonamiento. |
| `qwen3.8-flash` · `mimo-v2.5` · `mimo-v2.6-flash` | aceptado, profundidad no ajustable | El parámetro se acepta y nunca se rechaza, pero estos modelos gestionan su propia profundidad de razonamiento. |

Sin parámetro, cada modelo usa su propio valor por defecto (razonamiento
activo en `qwen3.6` y `gemma4`, con presupuesto de 16.384 tokens). Más
Expand Down
7 changes: 6 additions & 1 deletion src/content/docs-es/opencode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,11 @@ Escribe esto en `~/.config/opencode/opencode.json` para tenerlo en todos tus pro
"limit": { "context": 1048576, "output": 32768 },
"modalities": { "input": ["text", "image", "audio"], "output": ["text"] }
},
"mimo-v2.6-flash": {
"name": "Xiaomi MiMo V2.6 Flash",
"limit": { "context": 1048576, "output": 32768 },
"modalities": { "input": ["text", "image", "audio"], "output": ["text"] }
},
"gemma4": {
"name": "Gemma 4",
"limit": { "context": 262144, "output": 65536 },
Expand All @@ -76,7 +81,7 @@ Escribe esto en `~/.config/opencode/opencode.json` para tenerlo en todos tus pro
}
```

Esta es la configuración de los 7 modelos LLM que sirve NaN: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `gemma4`, `qwen3.6` y `glm5.3`.
Esta es la configuración de los 8 modelos LLM que sirve NaN: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` y `glm5.3`.

`@ai-sdk/openai-compatible` es el adaptador genérico, el que habla con cualquier API con forma de OpenAI. No uses `@ai-sdk/openai` a secas: ese espera la API de OpenAI de verdad.

Expand Down
12 changes: 10 additions & 2 deletions src/content/docs-es/pi.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,14 @@ En `~/.pi/agent/models.json`:
"contextWindow": 1048576,
"maxTokens": 32768
},
{
"id": "mimo-v2.6-flash",
"name": "Xiaomi MiMo V2.6 Flash",
"reasoning": true,
"input": ["text", "image"],
"contextWindow": 1048576,
"maxTokens": 32768
},
{
"id": "gemma4",
"name": "Gemma 4",
Expand Down Expand Up @@ -98,7 +106,7 @@ En `~/.pi/agent/models.json`:
}
```

Esta es la configuración de los 7 modelos LLM que sirve NaN: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `gemma4`, `qwen3.6` y `glm5.3`. Es el mismo conjunto que escribe el [CLI de NaN](/es/docs/nan-cli), y el mismo que publica la página de [OpenCode](/es/docs/opencode).
Esta es la configuración de los 8 modelos LLM que sirve NaN: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` y `glm5.3`. Es el mismo conjunto que escribe el [CLI de NaN](/es/docs/nan-cli), y el mismo que publica la página de [OpenCode](/es/docs/opencode).

`api: "openai-completions"` es lo que le dice a Pi con qué formato hablar. `maxTokens` es el techo de una respuesta, no el contexto. Es la misma cifra que el bloque de [OpenCode](/es/docs/opencode) publica como `limit.output`, porque responde a la misma pregunta.

Expand Down Expand Up @@ -136,7 +144,7 @@ Pídele algo corto. Si contesta, está saliendo por el clúster. Puedes cambiar
- **La clave va escrita en el fichero.** `models.json` vive en tu carpeta personal, así que no suele acabar en un repositorio, pero tenlo en cuenta si sincronizas tu configuración entre máquinas.
- **`maxTokens` no es el contexto.** Es el techo de cada respuesta. El razonamiento sale de ese mismo presupuesto, así que si pides respuestas largas y razonadas, súbelo.
- **Los modelos que declares son los que verás.** Pi no le pregunta al clúster qué hay disponible: muestra lo que haya en la lista.
- **`mimo-v2.5` entra sin su audio.** El esquema de Pi acepta `"text"` e `"image"` en `input` y nada más, y un tercer valor no tumba solo a ese modelo: Pi rechaza el fichero entero, con todos los demás proveedores que haya dentro. El modelo sigue oyendo audio por la API, pero no desde Pi.
- **Los dos MiMo entran sin su audio.** `mimo-v2.5` y `mimo-v2.6-flash`: el esquema de Pi acepta `"text"` e `"image"` en `input` y nada más, y un tercer valor no tumba solo a ese modelo: Pi rechaza el fichero entero, con todos los demás proveedores que haya dentro. Los modelos siguen oyendo audio por la API, pero no desde Pi.
- **`glm5.3` necesita una clave del tier premium.** Está en la lista para que quien lo tenga lo vea. Sin ese tier la respuesta es un `401`, explicado en [Elige tu modelo](/es/docs/choose-a-model).

</Details>
9 changes: 9 additions & 0 deletions src/content/docs-es/vscode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,15 @@ Se abre un fichero `chatLanguageModels.json`. Déjalo así:
"maxInputTokens": 1015808,
"maxOutputTokens": 32768
},
{
"id": "mimo-v2.6-flash",
"name": "Xiaomi MiMo V2.6 Flash",
"url": "https://api.nan.builders/v1/chat/completions",
"toolCalling": true,
"vision": true,
"maxInputTokens": 1015808,
"maxOutputTokens": 32768
},
{
"id": "gemma4",
"name": "Gemma 4",
Expand Down
3 changes: 2 additions & 1 deletion src/content/docs/choose-a-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ If you do not know which one to pick, look for what you want to do in the first
| Drive a coding agent through long sessions | `glm5.3` | It is built for that. Needs the premium tier |
| The same, but without the premium tier | `glm5.3-flash` | Same 1M context and a generous quota |
| Get an answer fast | `qwen3.8-flash` | Less depth, much less waiting |
| Hand the model an audio file directly | `mimo-v2.5` | It is the only one that hears |
| Hand the model an audio file directly | `mimo-v2.5` or `mimo-v2.6-flash` | Both hear audio natively |
| Describe or analyze an image | `deepseek-v4-flash` | Any of them except `glm5.3` will do; this is the best |
| Try things without spending quota | `gemma4` | It has no token counter |
| Build a search engine or a RAG | `qwen3-embedding` and then `rerank` | First you retrieve by similarity, then you reorder by relevance |
Expand All @@ -45,6 +45,7 @@ If you do not know which one to pick, look for what you want to do in the first
| `glm5.3-flash` | Coding agents, without premium | 1M | text · image | 2B tokens/month |
| `qwen3.8-flash` | Fast answers | 262K | text · image | 500M tokens/month |
| `mimo-v2.5` | Audio input, omnimodal | 1M | text · image · audio | 1.0B tokens/month |
| `mimo-v2.6-flash` | The newest MiMo, omnimodal | 1M | text · image · audio | 1.0B tokens/month |
| `gemma4` | Short tasks and testing | 262K | text · image | no counter |
| `qwen3.6` | Previous generation | 262K | text · image | no counter |
| `qwen3-embedding` | 4096-dimension vectors | - | text | no counter |
Expand Down
25 changes: 24 additions & 1 deletion src/content/docs/models.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,29 @@ with the same `base URL`.
]}
/>

<ModelCard
id="mimo-v2-6-flash"
name="mimo-v2.6-flash"
tag="omnimodal"
leftLabel="omnimodal — text, vision & audio"
rightLabel="capabilities"
description="The newest Xiaomi MiMo, natively omnimodal with vision and audio input. 1M token context. Tool calling and reasoning. Same limits as mimo-v2.5: 1.0B token monthly quota per member."
specs={[
{ label: 'Context', value: '1M tokens' },
{ label: 'Input modalities', value: 'text · image · audio' },
{ label: 'Output modalities', value: 'text' },
{ label: 'Monthly quota', value: '1.0B tokens / member' },
]}
items={[
'Tool calling (function calling)',
'Reasoning mode (recommended <code>max_tokens &ge; 300</code>)',
'Vision (image input)',
'Audio (audio input)',
'1M token context',
'Streaming generation (SSE)',
]}
/>

<ModelCard
id="gemma4"
name="gemma4"
Expand Down Expand Up @@ -322,7 +345,7 @@ cannot apply is never an error.
| `qwen3.6` | `none`, `minimal`, `low`, `medium`, `high`, `max` | `none` and `minimal` skip the reasoning phase entirely. The other four cap it: low 2,048, medium 8,192, high 16,384, max 32,768 tokens. |
| `gemma4` | `none`, `minimal`, `low`, `medium`, `high`, `max` | Same as `qwen3.6`: off, or a reasoning budget between 2,048 and 32,768 tokens. |
| `deepseek-v4-flash` | any value (no effect) | The model decides per request how much to reason; the parameter never changes that. |
| `qwen3.8-flash` · `mimo-v2.5` | accepted, depth not adjustable | The parameter is accepted and never rejected, but these models manage their own reasoning depth. |
| `qwen3.8-flash` · `mimo-v2.5` · `mimo-v2.6-flash` | accepted, depth not adjustable | The parameter is accepted and never rejected, but these models manage their own reasoning depth. |

With no parameter, every model uses its own default (reasoning on for
`qwen3.6` and `gemma4`, with a 16,384-token budget). More reasoning costs
Expand Down
7 changes: 6 additions & 1 deletion src/content/docs/opencode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,11 @@ Write this into `~/.config/opencode/opencode.json` to have it in every project,
"limit": { "context": 1048576, "output": 32768 },
"modalities": { "input": ["text", "image", "audio"], "output": ["text"] }
},
"mimo-v2.6-flash": {
"name": "Xiaomi MiMo V2.6 Flash",
"limit": { "context": 1048576, "output": 32768 },
"modalities": { "input": ["text", "image", "audio"], "output": ["text"] }
},
"gemma4": {
"name": "Gemma 4",
"limit": { "context": 262144, "output": 65536 },
Expand All @@ -76,7 +81,7 @@ Write this into `~/.config/opencode/opencode.json` to have it in every project,
}
```

This is the config for the 7 LLM models NaN serves: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `gemma4`, `qwen3.6` and `glm5.3`.
This is the config for the 8 LLM models NaN serves: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` and `glm5.3`.

`@ai-sdk/openai-compatible` is the generic adapter, the one that speaks to any API shaped like OpenAI's. Do not use plain `@ai-sdk/openai`: that one expects the real OpenAI API.

Expand Down
12 changes: 10 additions & 2 deletions src/content/docs/pi.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,14 @@ In `~/.pi/agent/models.json`:
"contextWindow": 1048576,
"maxTokens": 32768
},
{
"id": "mimo-v2.6-flash",
"name": "Xiaomi MiMo V2.6 Flash",
"reasoning": true,
"input": ["text", "image"],
"contextWindow": 1048576,
"maxTokens": 32768
},
{
"id": "gemma4",
"name": "Gemma 4",
Expand Down Expand Up @@ -98,7 +106,7 @@ In `~/.pi/agent/models.json`:
}
```

This is the config for the 7 LLM models NaN serves: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `gemma4`, `qwen3.6` and `glm5.3`. It is the same set the [NaN CLI](/docs/nan-cli) writes, and the same one the [OpenCode](/docs/opencode) page publishes.
This is the config for the 8 LLM models NaN serves: `deepseek-v4-flash`, `glm5.3-flash`, `qwen3.8-flash`, `mimo-v2.5`, `mimo-v2.6-flash`, `gemma4`, `qwen3.6` and `glm5.3`. It is the same set the [NaN CLI](/docs/nan-cli) writes, and the same one the [OpenCode](/docs/opencode) page publishes.

`api: "openai-completions"` is what tells Pi which format to speak. `maxTokens` is the ceiling for one answer, not the context. It is the same figure the [OpenCode](/docs/opencode) block publishes as `limit.output`, because it answers the same question.

Expand Down Expand Up @@ -136,7 +144,7 @@ Ask it for something short. If it answers, it is going out through the cluster.
- **The key is written in the file.** `models.json` lives in your home directory, so it does not usually end up in a repository, but keep it in mind if you sync your configuration between machines.
- **`maxTokens` is not the context.** It is the ceiling for each answer. Reasoning comes out of that same budget, so if you ask for long reasoned answers, raise it.
- **The models you declare are the ones you get.** Pi does not ask the cluster what is available: it shows whatever is on the list.
- **`mimo-v2.5` goes in without its audio.** Pi's schema accepts `"text"` and `"image"` for `input` and nothing else, and a third value does not fail that one model: Pi refuses the whole file, every other provider in it included. The model still hears audio through the API, just not from inside Pi.
- **Both MiMo models go in without their audio.** `mimo-v2.5` and `mimo-v2.6-flash`: Pi's schema accepts `"text"` and `"image"` for `input` and nothing else, and a third value does not fail that one model: Pi refuses the whole file, every other provider in it included. The models still hear audio through the API, just not from inside Pi.
- **`glm5.3` needs a key on the premium tier.** It is on the list so that a member who has it can see it. Without that tier the answer is a `401`, which is explained in [Choose your model](/docs/choose-a-model).

</Details>
9 changes: 9 additions & 0 deletions src/content/docs/vscode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,15 @@ A `chatLanguageModels.json` file opens. Leave it like this:
"maxInputTokens": 1015808,
"maxOutputTokens": 32768
},
{
"id": "mimo-v2.6-flash",
"name": "Xiaomi MiMo V2.6 Flash",
"url": "https://api.nan.builders/v1/chat/completions",
"toolCalling": true,
"vision": true,
"maxInputTokens": 1015808,
"maxOutputTokens": 32768
},
{
"id": "gemma4",
"name": "Gemma 4",
Expand Down
7 changes: 7 additions & 0 deletions src/data/modelos.json
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,13 @@
"frontier": true,
"mostUsed": true
},
{
"id": "mimo-v2.6-flash",
"by": "Xiaomi",
"specs": "omnimodal · 1M context · vision · audio · tool calling · reasoning",
"cuota": "1.0B tokens/mes",
"frontier": true
},
{
"id": "glm5.3",
"by": "Z.ai",
Expand Down
Loading
Loading