Skip to content

SMOODEV-3342: operator defaults to gpt-6-luna, groq judge - #570

Merged
brentrager merged 1 commit into
mainfrom
SMOODEV-3342-models
Sep 27, 2026
Merged

brentrager merged 1 commit into
mainfrom
SMOODEV-3342-models

Conversation

@brentrager

Copy link
Copy Markdown
Contributor

Problem

The Smoo AI gateway's policy (2026-09-26) is that every default LLM call uses the gpt-6-luna family or a Groq alias. Every smooth-operator server still defaulted to claude-haiku-4-5, for both the main turn and the conversation-workflow judge.

The nightly eval scorecard was also mislabelled. It graded claude-haiku-4-5 and claude-sonnet-4-5, but this week's gateway spend logs for the smooth-operator-gha key show those requests being served by LiteLLM fallbacks: gemini-2.5-flash, gpt-4.1-mini, gpt-4.1 and gpt-6-sol. The Anthropic calls were failing (credits exhausted) and LiteLLM fell back without any error. So the scorecard rows named models it never measured.

Before / after

Before After
Main-turn default (Rust server, Lambda, dev-support, Go, Python, .NET, TS) claude-haiku-4-5 gpt-6-luna
Workflow judge default same as the main model (claude-haiku-4-5) its own default, groq-gpt-oss-120b (SMOOTH_AGENT_JUDGE_MODEL still overrides)
Judge max_tokens Rust 16; Go/Python/.NET/TS 200 512 in every port
Eval CHEAP_MODEL (default agent and judge) claude-haiku-4-5 gpt-6-luna
Nightly eval matrix haiku judged by sonnet; sonnet judged by sonnet groq-gpt-oss-120b and gpt-6-luna, both judged by gpt-6-luna
Live E2E tests (Rust, Go, Python, .NET, TS) claude-haiku-4-5 gpt-6-luna
Helm values.yaml, Argo CD example, SST example claude-haiku-4-5 gpt-6-luna
examples/.env.example gemini-2.5-flash gpt-6-luna
scripts/operator-serve.sh deepseek-v4-flash gpt-6-luna
GET /admin/settings for an unsaved org (Go/Python/.NET/TS) claude-haiku-4-5 gpt-6-luna

Details

  • The Go and Python servers never actually sent their own default. They left the model empty, so the engine chose its built-in claude-haiku-4-5. Go only used its constant for a span attribute. Both servers now pass the default to the engine explicitly. A new parity test in each checks that the wire request carries gpt-6-luna.
  • Why the judge cap had to go up. groq-gpt-oss-120b is a reasoning model, and reasoning tokens count against max_tokens. With a small cap the model spends everything on reasoning and returns no text. The judge then has no verdict and the workflow never advances. gpt-oss on Groq ran out even at 200 (SMOODEV-2427), and gpt-6-luna-fast returns status=incomplete at 16. The new cap of 512 matches the voice path. The prompt still asks for a one-word JSON verdict.
  • Refactors in Rust and .NET. Rust's judge LlmConfig construction is now a small pure function (judge_llm_config) so it can be tested. .NET gets ServerEnv.ResolveModel / ResolveJudgeModel. The host now also reads the canonical SMOOTH_AGENT_JUDGE_MODEL, with SMOOTH_JUDGE_MODEL kept as an alias.
  • Changeset: a minor bump on the lockstep anchor, because a default changed.

Left alone, and why

  • Opaque test fixtures. Examples: ServerConfig { model: "claude-haiku-4-5", judge_model: … } in mock-backed Rust integration tests, admin tests that store claude-sonnet-4-5 / claude-opus-4-8 as an org's model, model-ceiling and /model/info payloads, and max-tokens limit tables. These strings are round-tripped against mocks and send no traffic.
  • Pricing tables in the engine cores (DEFAULT_PRICING lives in the separate smooth-operator-core repo). It has no gpt-6-luna entry. Two tests that needed a locally priced turn now pin claude-haiku-4-5 as the fixture model: Go TestStreamingTurnEmitsGenAISpans and Python test_absent_and_all_zero_headers_are_both_unmeasured. In production the gateway's cost header is what gets used, so this does not change reported cost.
  • Rust temperature measurement table and historical comments. These record what was measured, not defaults.
  • examples/.env.example's OpenAI line (gpt-4o-mini). This is a direct-OpenAI example for a BYO gateway, not a Smoo AI gateway default.
  • Existing parity gaps, not widened here. The Go host still does not read SMOOTH_AGENT_MODEL / SMOOTH_AGENT_JUDGE_MODEL. Rust reports model: null for an unsaved org's settings, while the other ports report the default.

Verification

  • Rust: cargo fmt --check, plus clippy (-D warnings) and cargo test on the crates this PR touches: smooai-smooth-operator-server, -evals, -lambda, -example-dev-support, smooai-smooth-operator.
  • Go: go test and go vet for go/server (includes the testcontainers Postgres tests), plus gofmt.
  • Python: uv run pytest in python/server (414 passed) and python/ (72 passed, 1 live test skipped), plus ruff check and ruff format --check.
  • .NET: dotnet build (Release, 0 warnings) and dotnet test on the solution. All green.
  • TS: pnpm typecheck and vitest run in typescript/server: 43 files and 398 tests passed.
  • Live tests that need gateway credentials skip without them, as designed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CfZaWemmpghhtofBauYdti

The Smoo AI gateway's policy is that every default LLM call uses the
gpt-6-luna family or a Groq alias. Every smooth-operator server still
defaulted its main turn AND its workflow judge to claude-haiku-4-5. On the
gateway, those calls had been failing over silently to whatever LiteLLM's
fallback chain picked.

- Main-turn default claude-haiku-4-5 -> gpt-6-luna in the Rust server,
  Lambda and dev-support example, and in the Go, Python, .NET and TS servers,
  plus the Helm values, Argo CD and SST examples, examples/.env.example (which
  shipped gemini-2.5-flash) and scripts/operator-serve.sh. The Go and Python
  servers left the model empty, so the ENGINE's built-in claude-haiku-4-5 went
  on the wire regardless of their constants. They now pass the default
  explicitly.
- The judge gets its own default, groq-gpt-oss-120b, instead of following the
  main model. SMOOTH_AGENT_JUDGE_MODEL still overrides it; the .NET host now
  reads that canonical name too.
- Judge max_tokens -> 512 in every port (Rust was 16, the rest 200).
  gpt-oss is a reasoning model and reasoning counts against the cap. A small
  cap came back with no text, so there was no verdict and the workflow never
  advanced.
- Evals: CHEAP_MODEL -> gpt-6-luna. The nightly matrix now grades
  groq-gpt-oss-120b and gpt-6-luna, both judged by gpt-6-luna. The old
  claude-haiku/sonnet rows were really graded on gateway fallbacks
  (gemini-2.5-flash, gpt-4.1-mini, gpt-4.1, gpt-6-sol) while Anthropic was
  failing, so the scorecard named models it never measured. Live E2E tests
  now run on gpt-6-luna.

Parity tests in every language assert the new defaults and the judge cap.
Opaque model strings in mock-backed fixtures, and the engine pricing tables,
are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CfZaWemmpghhtofBauYdti
@changeset-bot

changeset-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 1154071

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 2 packages
Name Type
@smooai/smooth-operator Minor
@smooai/smooth-operator-web-chat-example Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@brentrager
brentrager merged commit df9778c into main Sep 27, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant