SMOODEV-3342: operator defaults to gpt-6-luna, groq judge - #570
Merged
Merged
Conversation
The Smoo AI gateway's policy is that every default LLM call uses the gpt-6-luna family or a Groq alias. Every smooth-operator server still defaulted its main turn AND its workflow judge to claude-haiku-4-5. On the gateway, those calls had been failing over silently to whatever LiteLLM's fallback chain picked. - Main-turn default claude-haiku-4-5 -> gpt-6-luna in the Rust server, Lambda and dev-support example, and in the Go, Python, .NET and TS servers, plus the Helm values, Argo CD and SST examples, examples/.env.example (which shipped gemini-2.5-flash) and scripts/operator-serve.sh. The Go and Python servers left the model empty, so the ENGINE's built-in claude-haiku-4-5 went on the wire regardless of their constants. They now pass the default explicitly. - The judge gets its own default, groq-gpt-oss-120b, instead of following the main model. SMOOTH_AGENT_JUDGE_MODEL still overrides it; the .NET host now reads that canonical name too. - Judge max_tokens -> 512 in every port (Rust was 16, the rest 200). gpt-oss is a reasoning model and reasoning counts against the cap. A small cap came back with no text, so there was no verdict and the workflow never advanced. - Evals: CHEAP_MODEL -> gpt-6-luna. The nightly matrix now grades groq-gpt-oss-120b and gpt-6-luna, both judged by gpt-6-luna. The old claude-haiku/sonnet rows were really graded on gateway fallbacks (gemini-2.5-flash, gpt-4.1-mini, gpt-4.1, gpt-6-sol) while Anthropic was failing, so the scorecard named models it never measured. Live E2E tests now run on gpt-6-luna. Parity tests in every language assert the new defaults and the judge cap. Opaque model strings in mock-backed fixtures, and the engine pricing tables, are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CfZaWemmpghhtofBauYdti
🦋 Changeset detectedLatest commit: 1154071 The changes in this PR will be included in the next version bump. This PR includes changesets to release 2 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The Smoo AI gateway's policy (2026-09-26) is that every default LLM call uses the
gpt-6-lunafamily or a Groq alias. Every smooth-operator server still defaulted toclaude-haiku-4-5, for both the main turn and the conversation-workflow judge.The nightly eval scorecard was also mislabelled. It graded
claude-haiku-4-5andclaude-sonnet-4-5, but this week's gateway spend logs for thesmooth-operator-ghakey show those requests being served by LiteLLM fallbacks:gemini-2.5-flash,gpt-4.1-mini,gpt-4.1andgpt-6-sol. The Anthropic calls were failing (credits exhausted) and LiteLLM fell back without any error. So the scorecard rows named models it never measured.Before / after
claude-haiku-4-5gpt-6-lunaclaude-haiku-4-5)groq-gpt-oss-120b(SMOOTH_AGENT_JUDGE_MODELstill overrides)max_tokensCHEAP_MODEL(default agent and judge)claude-haiku-4-5gpt-6-lunagroq-gpt-oss-120bandgpt-6-luna, both judged bygpt-6-lunaclaude-haiku-4-5gpt-6-lunavalues.yaml, Argo CD example, SST exampleclaude-haiku-4-5gpt-6-lunaexamples/.env.examplegemini-2.5-flashgpt-6-lunascripts/operator-serve.shdeepseek-v4-flashgpt-6-lunaGET /admin/settingsfor an unsaved org (Go/Python/.NET/TS)claude-haiku-4-5gpt-6-lunaDetails
claude-haiku-4-5. Go only used its constant for a span attribute. Both servers now pass the default to the engine explicitly. A new parity test in each checks that the wire request carriesgpt-6-luna.groq-gpt-oss-120bis a reasoning model, and reasoning tokens count againstmax_tokens. With a small cap the model spends everything on reasoning and returns no text. The judge then has no verdict and the workflow never advances. gpt-oss on Groq ran out even at 200 (SMOODEV-2427), andgpt-6-luna-fastreturnsstatus=incompleteat 16. The new cap of 512 matches the voice path. The prompt still asks for a one-word JSON verdict.LlmConfigconstruction is now a small pure function (judge_llm_config) so it can be tested. .NET getsServerEnv.ResolveModel/ResolveJudgeModel. The host now also reads the canonicalSMOOTH_AGENT_JUDGE_MODEL, withSMOOTH_JUDGE_MODELkept as an alias.Left alone, and why
ServerConfig { model: "claude-haiku-4-5", judge_model: … }in mock-backed Rust integration tests, admin tests that storeclaude-sonnet-4-5/claude-opus-4-8as an org's model, model-ceiling and/model/infopayloads, and max-tokens limit tables. These strings are round-tripped against mocks and send no traffic.DEFAULT_PRICINGlives in the separate smooth-operator-core repo). It has nogpt-6-lunaentry. Two tests that needed a locally priced turn now pinclaude-haiku-4-5as the fixture model: GoTestStreamingTurnEmitsGenAISpansand Pythontest_absent_and_all_zero_headers_are_both_unmeasured. In production the gateway's cost header is what gets used, so this does not change reported cost.temperaturemeasurement table and historical comments. These record what was measured, not defaults.examples/.env.example's OpenAI line (gpt-4o-mini). This is a direct-OpenAI example for a BYO gateway, not a Smoo AI gateway default.SMOOTH_AGENT_MODEL/SMOOTH_AGENT_JUDGE_MODEL. Rust reportsmodel: nullfor an unsaved org's settings, while the other ports report the default.Verification
cargo fmt --check, plus clippy (-D warnings) andcargo teston the crates this PR touches:smooai-smooth-operator-server,-evals,-lambda,-example-dev-support,smooai-smooth-operator.go testandgo vetforgo/server(includes the testcontainers Postgres tests), plusgofmt.uv run pytestinpython/server(414 passed) andpython/(72 passed, 1 live test skipped), plusruff checkandruff format --check.dotnet build(Release, 0 warnings) anddotnet teston the solution. All green.pnpm typecheckandvitest runintypescript/server: 43 files and 398 tests passed.🤖 Generated with Claude Code
https://claude.ai/code/session_01CfZaWemmpghhtofBauYdti