Synthwave makes several AI models smarter by making them work together. It is a Mixture-of-Agents (MoA) self-hosted server by Trac Systems (you streamline multiple models into one single isntance): instead of routing your prompt to one model, Synthwave asks several independent models to each draft an answer, then has a final synthesizer model read every draft and fuse them into one best response. The combined system reasons better, makes fewer mistakes, and covers more ground than any of its individual members — all behind a single OpenAI-compatible endpoint.
You call it like any normal model API (POST /v1/chat/completions with model: "synthwave"). Operators define the model fleet and how the models combine — parallel fan-out with synthesis, fallback cascades, or voting — in a TOML config.
Important
Synthwave scores 90.0% (27/30) on AIME 2026 — landing in the frontier band, above DeepSeek V3.2 and Claude Opus 4.6, on a freshly released, uncontaminated competition set. This is achieved with internal reasoning mostly off.
Measured on the full AIME 2026 set (American Invitational Mathematics Examination I + II, 30 problems, single run; final answers graded as exact integers). The three misses are the three hardest items — the two final AIME II problems and one diagram-dependent problem.
| Model | AIME 2026 | Internal thinking |
|---|---|---|
| GPT-5 | ~100% | full |
| Gemini 3.1 Pro | ~95% | full |
| Grok 4.2 | ~93% | full |
| Synthwave | 90.0% | mostly off |
| DeepSeek V3.2 | ~88% | full |
| Claude Opus 4.6 | ~85% | full |
| Llama 4 Scout | ~80% | full |
| Qwen 3.5 | ~76% | full |
Competitor figures come from public AIME 2026 leaderboards and can be checked directly: MathArena — AIME 2026 and the LLM leaderboard. Only the Synthwave row is measured by us; treat the other rows as indicative (sources and exact evaluation conditions vary).
Note
Thinking was mostly off — and turning it on may increase intelligence. Of the four models in the ensemble, only one generator (Ornith) runs with internal step-by-step reasoning enabled; the other two generators and the synthesizer run without it. So 90% reflects an ensemble that is largely not using extended thinking. Enabling reasoning across more of the fleet is untested headroom that may push the score higher.
The configuration above runs across two NVIDIA DGX Spark (GB10) nodes, each model served with vLLM behind an OpenAI-compatible endpoint and NVFP4-quantized for the GB10 hardware. The Mixture-of-Agents profile fans out to three generators (each writes an independent draft, in parallel) and fuses their drafts with one synthesizer:
| Role | Model | Notes |
|---|---|---|
| Generator | Ornith-1.0-35B | reasoning generator — the only member with internal thinking on |
| Generator | Qwen3.6-35B-A3B | generalist (also the image/video input path) |
| Generator | Qwen3-Coder-Next | code-focused generator |
| Synthesizer | Gemma-4-26B-A4B | reads every draft, writes the final fused answer (thinking off) |
Synthesis settings: merge mode, generator temperature 0.6, synthesizer temperature 0.2. The generators are queried concurrently; the synthesizer then reads all drafts and produces a single answer. Only Ornith uses extended thinking — the other generators and the synthesizer do not.
This runs on two DGX Spark (GB10) desktop-class nodes — not a datacenter — yet the ensemble reaches frontier territory for intelligence. None of the individual models is a frontier model: three mid-size (26–35B) generators, each on its own well below the top of the table, are fused into a result that ranks alongside the best. The intelligence is a property of the composition, not of any single model or of raw scale.
That principle scales down, too. The Mixture-of-Agents mechanism is model-agnostic — it lifts the effective intelligence of whatever fleet it is given, including single-machine setups running smaller models. You do not need this exact lineup or this hardware; you need a diverse, balanced set of models and a synthesizer to combine them. Because intelligence here depends on the composition of models, improving the mix — stronger generators, more diversity, more reasoning enabled — raises the result further than scaling any one model could.
Important
On the hard split of LiveCodeBench — the benchmark's most difficult tier — a cheap, reasoning-off Synthwave ensemble scores 46.3% pass@1 (38/82) for ≈$2.66 the full run (≈$0.032/problem) — more than any of its own generators scores alone (glm-5.2 29.3%, qwen3-coder-next 35.4%, qwen3.5-122b 39.0%). gpt-5.5 alone scores 73.2% (60/82) but costs ≈$14.78 (≈$0.18/problem) — about 5.6× more.
Measured on the LiveCodeBench code-generation set, hard difficulty only (82 problems, release_latest, contests dated 2025-01-01 onward), single run, pass@1, official LiveCodeBench evaluation. Both configurations run with internal reasoning off. The contests postdate the models' training cutoff, and every generator draft is bounded so reasoning-in-content cannot pollute the synthesizer — the score is billed only on clean, parseable output, never on a truncated or reasoning-polluted draft.
| Role | Model | Served via |
|---|---|---|
| Generator | glm-5.2 | OpenRouter (z-ai/glm-5.2) |
| Generator | qwen3.5-122b | OpenRouter (qwen/qwen3.5-122b-a10b) |
| Generator | qwen3-coder-next | OpenRouter (qwen/qwen3-coder-next) |
| Synthesizer | gemma-4-26b | OpenRouter (google/gemma-4-26b-a4b-it) |
Three generators each draft an answer with reasoning off; the synthesizer reads all three bounded drafts and fuses the final solution. The comparison model, gpt-5.5 (reasoning_effort: medium), is run standalone on the OpenAI API.
| Setup | pass@1 (hard) | cost — full run | per problem |
|---|---|---|---|
| glm-5.2 — member, alone, reasoning off | 29.3% (24/82) | — | — |
| qwen3-coder-next — member, alone, reasoning off | 35.4% (29/82) | — | — |
| qwen3.5-122b — member, alone, reasoning off | 39.0% (32/82) | — | — |
| Synthwave — cloud ensemble, reasoning off | 46.3% (38/82) | ≈$2.66 | ≈$0.032 |
| gpt-5.5 — alone, reasoning medium | 73.2% (60/82) | ≈$14.78 | ≈$0.18 |
No single member matches the ensemble: run alone, glm-5.2 scores 29.3%, qwen3-coder-next 35.4%, and qwen3.5-122b 39.0% — all below the ensemble's 46.3%. The accuracy is a property of the composition, not of any one model. gpt-5.5 alone is stronger but costs ≈5.6× more to run; the ensemble recovers ≈2/3 of its hardest-split accuracy at ≈1/5 of the cost, from hosted commodity models with reasoning off.
Note
This is the hardest slice of LiveCodeBench — pass rates here run well below the full set. gpt-5.5 is higher and no single model at the ensemble's price point matches it; the ensemble's 46.3% is a property of the composition, not of any one member. But for a fraction of the cost it delivers most of the way there — the accuracy-per-dollar tradeoff that matters when budget is the constraint. Costs are measured: the Synthwave run on OpenRouter (full 82-problem run); gpt-5.5 from real per-request OpenAI usage sampled across 7 hard problems (×82) at list prices ($5/1M in, $30/1M out), June 2026. gpt-5.5 accuracy is a single run (a prior run scored 79.3% — reasoning models carry run-to-run variance).
Synthwave exposes a normal /v1 model API while internally routing each request through one of several operator-defined profiles: single-upstream passthrough, parallel fan-out with synthesis, fallback cascades, or voting. Clients can use it as a drop-in model endpoint; operators control the model fleet and profile behavior in TOML.
- OpenAI-compatible
POST /v1/chat/completions. - Compatibility adapters for
POST /v1/completionsandPOST /v1/responses. GET /v1/modelswith profile/model capabilities.GET /v1/healthreadiness checks against configured upstreams.GET /v1/metrics/moafor recent MoA dispatch telemetry.- Optional bearer auth on the public API.
- Per-upstream bearer or basic auth.
- Server-owned profiles for MoA, cascade, and voting behavior.
- Tool/function-call normalization and sanitization.
- Optional image, video, and audio routing through server-level cascades.
- Optional single-model facade via
server.model_name.
src/meta_model/ FastAPI service and dispatch code
src/meta_model/moa/ Fan-out, synthesis, multimodal, and tool policy
tests/ Contract and behavior tests
meta-model.toml.example Annotated operator configuration
pyproject.toml Python package metadata
PAPER.md System paper and design notes
- Python
3.11+. - One or more upstream model endpoints that speak an OpenAI-compatible
/v1/chat/completionsAPI. - Optional upstream
/v1/modelssupport for readiness probes.
The upstreams can be local inference servers, private network services, cloud gateways, or a mixture of those. Synthwave only needs the base URL, model id, context budget, output budget, modalities, and auth settings.
cd synthwave
python3.11 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"For a runtime-only install, omit the dev extra:
pip install -e .Start from the example config:
cp meta-model.toml.example meta-model.toml
$EDITOR meta-model.tomlThe config has three layers:
server: public API binding, auth, CORS, facade name, and optional system identity.upstreams: available models and how to call them.profiles: callable model/profile names exposed through/v1/models.
Use this when you want Synthwave to expose one existing model as an OpenAI-compatible endpoint with health checks, auth, and a stable model name.
[server]
host = "127.0.0.1"
port = 8400
bearer_token = "change-me"
model_name = "synthwave-local"
[upstreams.main]
model_id = "provider/model-id"
base_url = "http://127.0.0.1:8000/v1"
context = 32768
max_output = 4096
modalities = ["text"]
api_key_env = "MAIN_MODEL_KEY"
[profiles."chat.v1"]
type = "cascade"
upstreams = ["main"]
aliases = ["synthwave-local"]Run with:
export MAIN_MODEL_KEY="upstream-token-if-needed"
export META_MODEL_CONFIG=./meta-model.toml
meta-model --config "$META_MODEL_CONFIG"Or directly with Uvicorn:
META_MODEL_CONFIG=./meta-model.toml \
uvicorn meta_model.server:app --host 127.0.0.1 --port 8400Use this when you have multiple models and want independent drafts synthesized into one final answer.
[server]
host = "127.0.0.1"
port = 8400
bearer_token = "change-me"
model_name = "synthwave"
request_timeout_secs = 600
[upstreams.primary]
model_id = "your-best-synthesizer"
base_url = "http://127.0.0.1:8001/v1"
context = 131072
max_output = 8192
modalities = ["text"]
api_key_env = "PRIMARY_MODEL_KEY"
[upstreams.fast]
model_id = "your-fast-coder"
base_url = "http://127.0.0.1:8002/v1"
context = 32768
max_output = 4096
modalities = ["text"]
api_key_env = "FAST_MODEL_KEY"
[upstreams.reasoning]
model_id = "your-reasoning-model"
base_url = "http://127.0.0.1:8003/v1"
context = 32768
max_output = 4096
modalities = ["text"]
api_key_env = "REASONING_MODEL_KEY"
request_overrides = { include_reasoning = false }
[profiles."write_synth.v1"]
type = "moa"
generators = ["fast", "reasoning"]
synthesizer = "primary"
synthesis_mode = "merge"
generator_temperature = 0.4
synthesizer_temperature = 0.2
non_client_synth_reserve_tokens = 8192
aliases = ["synthwave"]How to choose upstream roles:
synthesizer: use the model with the best instruction following, longest context, and strongest final-answer quality.generators: use diverse models. Speed matters because the MoA wall clock is bounded by the slowest required generator.context: set the real usable input context, not the marketing maximum.max_output: set the largest generation budget you are willing to allocate to that upstream.modalities: use['text']for text-only models and addimage,video, oraudioonly when the upstream really supports that input shape.request_overrides: use for server-owned upstream quirks such as disabling hidden reasoning output, forcing a parser mode, or setting model-specific request fields.
Opt-in profile flag, default false. Set it when the profile backs an autonomous agent client — opencode, an Agent SDK loop, or any client that drives a multi-turn tool loop and only keeps going while the model returns tool_calls.
[profiles."agent.v1"]
type = "moa"
generators = ["fast", "reasoning"]
synthesizer = "primary"
synthesis_mode = "merge"
agentic_synth = trueBy default, under tool_choice = "auto" the merge synthesizer is handed the request's tools only when at least one generator already emitted a tool_call that turn. If every generator answered in plain text — narrating progress or asking what to do next — the synthesizer has no tools to work with and returns text, which makes the agent client stop the loop and surface a "what next?" message instead of advancing.
With agentic_synth = true the synthesizer keeps the tools under auto even when no generator called one, and runs an agent-continuation system prompt. It emits the next tool_call when the task is still in progress, answers in text only when the task is genuinely finished or truly blocked on the user, and never claims it performed an action (wrote a file, ran a command) without actually emitting the matching tool_call.
It does not force a tool call — auto stays optional, so plain chat still gets a text answer. Leave it false for general chat or synthesis profiles; the text-merge path is then byte-identical to before. It has no effect on best-of profiles, whose synthesizer returns an integer index rather than authoring tool calls.
Opt-in profile flag, default false (or set globally with the TK_MOA_GENERATOR_RATIONALE env var, mirroring strict_quorum / TK_MOA_STRICT_QUORUM). It hardens the drafts the synthesizer reads.
[profiles."moa.v1"]
type = "moa"
generators = ["gen_a", "gen_b"]
synthesizer = "judge"
synthesis_mode = "merge"
generator_rationale = trueModels run with internal reasoning nominally off often reroute their chain-of-thought into the visible answer instead of a separate reasoning channel. Forwarding those raw drafts to the synthesizer bloats the synthesis context with tens of thousands of tokens of prose and can bury the actual answer. With generator_rationale = true, each generator is instructed to emit a short [RATIONALE] … [/RATIONALE] block followed by its final answer; the synthesizer then receives only (bounded rationale + answer), with any pre-block chain-of-thought stripped.
It is a contract, not a detector — instructing the model to fence its rationale is what makes untagged reasoning strippable (a plain heuristic can't tell reasoning from answer). The sentinel is the square-bracket [RATIONALE] form, not <rationale> (the angle form is treated as chain-of-thought and would be deleted). For generators that ignore the contract and dump unbounded reasoning, the strip salvages the final fenced code/answer block and discards the ramble, so a non-compliant member still can't pollute synthesis. Override the instruction text with generator_rationale_prompt; both take effect after a restart. Leave it false and the draft path is byte-identical to before.
Multimodal routing is server-level. Profiles do not opt in individually. Configure ranked cascades by upstream name:
[upstreams.vision_primary]
model_id = "your-vision-model"
base_url = "http://127.0.0.1:8010/v1"
context = 32768
max_output = 4096
modalities = ["text", "image"]
[vision]
endpoints = ["vision_primary"]
[video]
endpoints = []
[audio]
endpoints = []If a modality list is empty, requests using that input type return a typed unsupported-modality error instead of being sent to a text-only model.
Every upstream speaks a protocol. The default, "openai", is any OpenAI-compatible /v1/chat/completions endpoint (vLLM, the OpenAI API, a gateway). Opt-in extras let hosted cloud models join the same ensemble — as generators, the synthesizer, or vision endpoints — without touching existing configs (all default off, so a config with no provider blocks is byte-identical to before): "anthropic" (native Messages API), and "openai-responses" (OpenAI's Responses API, for reasoning models that need tools + effort together).
Set protocol = "anthropic" and the upstream is driven through synthwave's native Anthropic Messages (/v1/messages) adapter instead of /chat/completions. The adapter translates synthwave's internal OpenAI request/response shape to and from Anthropic's wire format, so the rest of the system (MoA fan-out, synthesis, cascade, vision routing, tool calling) treats the Claude model like any other upstream.
Why a native adapter and not Anthropic's OpenAI-compatibility endpoint? That shim accepts
reasoning_effortbut silently ignores it — you get no extended thinking. The native Messages path is the only way to actually control effort.
Minimal config:
[upstreams.opus]
model_id = "claude-opus-4-8"
base_url = "https://api.anthropic.com/v1" # adapter appends /messages
context = 200000
max_output = 32000
modalities = ["text", "image"] # Claude vision works through the adapter
supports_thinking = true # advertise thinking in /v1/models
protocol = "anthropic"
api_key_env = "ANTHROPIC_API_KEY" # sent as the x-api-key header
[upstreams.opus.anthropic]
version = "2023-06-01" # anthropic-version header (default shown)
thinking = "adaptive" # "adaptive" | "enabled" | "off"
effort = "xhigh" # adaptive ceiling → output_config.effort: low|medium|high|xhigh
# budget_tokens = 8000 # ONLY for thinking = "enabled" (legacy fixed-budget models)
# drop_params = ["temperature", "top_p"] # strip params newer models (opus-4-8) reject, even thinking-off
# cache = true # prompt caching: cache system+tools+history; cache reads ~0.1x inputThen export the key and run:
export ANTHROPIC_API_KEY="sk-ant-..."
META_MODEL_CONFIG=meta-model.toml meta-model[upstreams.opus.anthropic] options:
| Field | Meaning |
|---|---|
version |
anthropic-version request header. Default "2023-06-01". |
thinking |
"adaptive" (model decides how much to think, bounded by effort — the Claude 4.x knob), "enabled" (legacy fixed budget, needs budget_tokens), or "off"/omitted (no thinking). |
effort |
Reasoning ceiling for adaptive thinking → output_config.effort: low/medium/high/xhigh. A ceiling, not a floor — easy prompts spend ~0 thinking tokens, so cost stays low. |
budget_tokens |
Fixed thinking budget for thinking = "enabled" only. The adapter raises max_tokens above it automatically. |
drop_params |
Sampling params to strip before forwarding, e.g. ["temperature", "top_p"]. Anthropic's newest models (claude-opus-4-8) deprecate these and 400 if they are sent — even with thinking off. Only consulted when thinking is off (thinking-on already drops them). Default []. |
cache |
true enables Anthropic prompt caching: the adapter marks ephemeral cache_control breakpoints on the stable prefix — tools, system, and the conversation-so-far. Cache reads bill at ~0.1× input, so agent loops that resend a big system prompt + tool defs + growing history every turn get a large saving. GA feature, 5-minute ephemeral TTL, no beta header. Default false. |
What the adapter handles for you:
- System messages are lifted to Anthropic's top-level
system; the rest become alternating user/assistant turns (consecutive same-role turns are merged). - Tool calling — OpenAI
tools/tool_choice↔ Anthropictools/tool_use/tool_result, both directions, so agentic clients work unchanged. - Images — OpenAI
image_urlparts (bothdata:base64 URIs and remote URLs) become Anthropic image blocks. - Thinking & effort — sent as
thinking+output_config.effort; sampling params Anthropic rejects under thinking (temperature,top_p) are dropped automatically. - Auth — the standard
api_key/api_key_envfields are sent as thex-api-keyheader (no separate config). - Prompt caching (opt-in,
cache = true) — ephemeralcache_controlbreakpoints on tools + system + the conversation prefix; Anthropic reads the longest matching prefix from cache (~0.1× input) and only writes the delta. Cache reads/writes surface in the OpenAI-shaped response asusage.prompt_tokens_details.cached_tokens, so hits are visible in normal usage accounting.
Thinking-mode compatibility (Claude 4.x — set thinking to the mode the model supports; verified live):
| Model | "adaptive" (+ effort) |
"enabled" (+ budget_tokens) |
|---|---|---|
claude-opus-4-8 |
✅ | ❌ rejects |
claude-sonnet-4-6 |
✅ | ✅ |
claude-opus-4-5, claude-sonnet-4-5, claude-opus-4-1, claude-haiku-4-5 |
❌ rejects | ✅ |
Rule of thumb: the newest models take adaptive (+ effort); older / smaller 4.x take the legacy enabled (+ budget_tokens); sonnet-4-6 takes either. Any model also accepts thinking = "off" (or just omit the block). Everything else — tools, images, system handling, response shape — is model-agnostic across the whole family, so switching models is a one-line config change.
Other notes (verified live):
- Claude 4.x returns redacted thinking — an empty thinking block, no readable reasoning text. The effort signal is the token count, surfaced in the OpenAI response as
usage.completion_tokens_details.reasoning_tokens(matching OpenAI reasoning models). Soreasoning_contentwill normally be absent even though thinking occurred; checkreasoning_tokensto confirm effort engaged.
These stay on protocol = "openai" — same wire format synthwave already speaks, no adapter needed. The catch is that synthwave (like any MoA) injects a generator/synthesizer temperature and forwards the caller's max_tokens, and OpenAI's reasoning models reject both. Against the live API a stock request returns:
max_tokens -> 400 "Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead."
temperature -> 400 "Unsupported value: 'temperature' does not support 0.6 ... Only the default (1) value is supported."
So a reasoning model will not work on the default path unmodified. The optional [upstreams.<name>.openai] block normalizes the request so it does:
[upstreams.gpt]
model_id = "gpt-5.5"
base_url = "https://api.openai.com/v1" # adapter appends /chat/completions
context = 400000
max_output = 16000
modalities = ["text", "image"] # gpt-5.x vision passes through natively
api_key_env = "OPENAI_API_KEY" # sent as Authorization: Bearer <key>
[upstreams.gpt.openai]
reasoning_effort = "high" # none|minimal|low|medium|high|xhigh (subset varies by model)
max_tokens_param = "max_completion_tokens" # rename the caller's max_tokens
drop_params = ["temperature", "top_p"] # strip params the model 400s onThen export the key and run:
export OPENAI_API_KEY="sk-proj-..."
META_MODEL_CONFIG=meta-model.toml meta-model[upstreams.gpt.openai] options:
| Field | Meaning |
|---|---|
reasoning_effort |
Injected into the request: none/minimal/low/medium/high/xhigh. The accepted subset varies by model (the upstream validates); omit to use the model default. none/minimal run the model effectively thinking-off (zero reasoning tokens). |
max_tokens_param |
"max_completion_tokens" renames the caller's max_tokens to the field reasoning models require. Leave as the default "max_tokens" for normal OpenAI-compatible models. |
drop_params |
List of request fields to strip before forwarding. Use ["temperature", "top_p"] for reasoning models, which 400 on a non-default temperature. |
Model-class compatibility (which [openai] knobs each class needs; verified live):
| Model class (examples) | [openai] block needed |
|---|---|
Reasoning — gpt-5.5, gpt-5-mini, o3-mini, … |
Yes — rename + drop_params (these reject both max_tokens and a non-default temperature) |
Newer chat — gpt-5.3-chat-latest |
Yes — at least max_tokens_param (rejects max_tokens) |
Older chat — gpt-5-chat-latest |
No — accepts temperature + max_tokens, runs on the default path |
reasoning_effort levels vary by model — e.g. gpt-5-mini / gpt-5-nano accept minimal/low/medium/high (no xhigh), while the gpt-5.4 family adds none (a full thinking-off) and xhigh. The block forwards whatever you set and the upstream validates it, so match it to the model; none/minimal give zero reasoning tokens (fastest/cheapest).
Notes:
- Images pass through unchanged — gpt-5.x accepts OpenAI
image_urlparts natively, so just declaremodalities = ["text", "image"]; no translation happens. - Reasoning usage is reported by OpenAI under
usage.completion_tokens_details.reasoning_tokens(the same field the Anthropic adapter populates), so effort spend is visible the same way across providers. - Plain (non-reasoning) OpenAI-compatible models need no block at all — the default
protocol = "openai"with no[openai]table is the original byte-identical path. Only add the block for models that rejectmax_tokens/temperature.
There is one combination the /chat/completions path above cannot serve: a gpt-5.x reasoning model that must use function tools and a reasoning_effort level at the same time. The live API rejects it outright:
tools + reasoning_effort -> 400 "Function tools with reasoning_effort are not supported
for gpt-5.5 in /v1/chat/completions. Please use /v1/responses instead."
This bites every agentic client (e.g. opencode always sends tools): a gpt-5.x generator or synthesizer that must also honor a fixed thinking level (off or on) can only run through OpenAI's Responses API (/v1/responses). Set protocol = "openai-responses" and the upstream is driven through synthwave's native Responses adapter, which translates synthwave's internal OpenAI request/response to and from the Responses wire shape (messages → input, system → instructions, tools flattened, tool calls/results → function_call/function_call_output, response output items → message/tool_calls). The rest of the system treats it like any other upstream.
It reuses the same [openai] block — no new knobs to learn:
[upstreams.gpt]
model_id = "gpt-5.5"
base_url = "https://api.openai.com/v1" # adapter appends /responses
context = 400000
max_output = 16000
modalities = ["text"]
protocol = "openai-responses" # <- the only change vs the openai block above
api_key_env = "OPENAI_API_KEY"
[upstreams.gpt.openai]
reasoning_effort = "none" # THINKING OFF. flip to low|medium|high|xhigh for THINKING ON
drop_params = ["temperature", "top_p"] # reasoning models reject thesereasoning_effortis the on/off switch, mapped to Responses'reasoning.effort:"none"= thinking off (zero reasoning tokens),"low"|"medium"|"high"|"xhigh"= thinking on. To A/B off vs on, change this one value and restart — nothing else. (gpt-5.5 rejects"minimal"on this endpoint; the model validates.)drop_paramsworks exactly as on theopenaiblock.max_tokens_paramis ignored here — the adapter always emitsmax_output_tokens(clamped tomax_output), which the Responses API requires.- Reasoning usage is surfaced under
usage.completion_tokens_details.reasoning_tokens, same as everywhere else.
Use protocol = "openai" for a gpt-5.x model in a plain chat profile (no tools) and protocol = "openai-responses" when the same model is a generator/synthesizer in a tool-using / agentic profile.
Because every protocol normalizes to synthwave's internal OpenAI contract, you can list cloud and local upstreams together in a single profile:
[profiles."cloud-moa.v1"]
type = "moa"
generators = ["gpt", "opus", "sonnet"] # cross-provider diversity
synthesizer = "opus"
fastpath_on_agreement = trueSee the annotated [upstreams.gpt.openai] and [upstreams.opus.anthropic] examples in meta-model.toml.example.
With the console script:
META_MODEL_CONFIG=./meta-model.toml \
meta-model --config ./meta-model.tomlWith Uvicorn:
META_MODEL_CONFIG=./meta-model.toml \
uvicorn meta_model.server:app --host 127.0.0.1 --port 8400If your config sets [server] host and port, prefer the console script so the service reads those values directly.
Health:
curl -s http://127.0.0.1:8400/v1/health | jqModel catalog:
curl -s http://127.0.0.1:8400/v1/models \
-H 'Authorization: Bearer change-me' | jqChat completion:
curl -s http://127.0.0.1:8400/v1/chat/completions \
-H 'Authorization: Bearer change-me' \
-H 'Content-Type: application/json' \
-d '{
"model": "synthwave",
"messages": [{"role": "user", "content": "Write a compact hello-world in Python."}],
"max_tokens": 200,
"temperature": 0.2
}' | jqSelect a specific profile:
{
"model": "write_synth.v1",
"x_meta_model": {"profile": "write_synth.v1"},
"messages": [{"role": "user", "content": "Implement the requested file."}],
"max_tokens": 1200
}GET /v1/health Readiness; 503 when any required upstream is unhealthy
GET /health Alias for /v1/health
GET /v1/models OpenAI-style model/profile catalog
POST /v1/chat/completions Main OpenAI-compatible chat endpoint
POST /v1/completions Legacy text-completion adapter
POST /v1/responses Responses-style adapter for supported request shapes
POST /tokenize Token counting via configured tokenizer upstream
GET /v1/metrics/moa Recent MoA dispatch telemetry
GET /metrics Alias for /v1/metrics/moa
- Treat
/v1/healthas readiness, not liveness. A model outage should stop traffic, not necessarily restart the service. - MoA latency is determined by the slowest required generator plus synthesis time.
- Keep
request_timeout_secsabove the expected worst-case fan-out and synthesis wall clock. - Use aliases or
server.model_namewhen external clients need one stable model id. - Keep secrets in environment variables with
api_key_envorbasic_auth_pass_env. - Use
GET /v1/metrics/moato inspect generator success, fallback behavior, draft sizes, and elapsed time.
cd synthwave
source .venv/bin/activate
pytest
ruff check .Synthwave Meta-Model is authored by Markus Bopp from Trac Systems.