Skip to content

Refuse unknown config keys and trainer flags the target does not accept - #162

Merged
ProfSynapse merged 3 commits into
feat/submodule-cloud-api-v1from
claude/config-strictness
Oct 2, 2026
Merged

ProfSynapse merged 3 commits into
feat/submodule-cloud-api-v1from
claude/config-strictness

Conversation

@ProfSynapse

Copy link
Copy Markdown
Owner

Summary

Config typos and unsupported settings used to change a training run without any error. They now stop the run.

This is PR 1 of 5 in a stack. Merge in order.

Unknown config keys

  • The SFT, KTO, DPO and GRPO loaders now refuse keys their schema doesn't declare, at any depth. The error lists each dotted path with a "did you mean" suggestion.
    • The shared helper is in shared/training_utils.py.
    • GRPO and env-GRPO get declared schemas next to the code that reads them. A test checks those schemas against the keys the code actually reads, so they can't drift apart.
  • Tier presets also refuse keys they don't route. Numeric strings such as 5e-4 are converted through one shared conversion function.
  • The Modal recipe _section errors now name the offending fields.

Settings dropped because the installed TRL doesn't accept them

  • GRPO no longer filters GRPOConfig arguments silently. Any argument from YAML, extra_args, use_gspo or the CLI that the installed TRL doesn't accept now raises an error. It names the argument and the installed trl version.
  • The one version-dependent internal default, max_prompt_length, is listed explicitly in code.

Launcher and trainer flag mismatches (found by a new contract test)

  • HF Jobs: passed --save-steps/--save-total-limit to KTO and DPO, and --seed to env-GRPO, which none of them accepted. Now wired.
  • HF Jobs: passed run flags to the embedding and ace_step trainers, which don't accept them. HF Jobs now refuses those methods.
  • Local-run: passed --model-revision to KTO and DPO, which have no revision pin. Local-run now refuses it.
  • Flywheel experiment loop: passed its override YAML as --config, which train_sft executed as Python. Hyperparameters now map to explicit trainer flags, and unmapped keys are refused.
  • Contract test: tests/contract/test_trainer_argv_contract.py builds every launcher's argv (HF Jobs, RunPod, local-run, RTX, Mac, flywheel, checked-in recipes, runtime_v1, protected smoke) and parses it with the trainer's real argument parser.

Dead config removed, or wired up

  • Removed:
    • GRPO's rewards.items[].params and rewards.custom.module, which were never read.
    • Env-GRPO's env_training.backend; the code reads env_backend.
    • Env-GRPO's --max-seq-length.
    • KTO/DPO dataset.chat_template.
  • Now wired: group_by_length, wandb.project and wandb.entity (all methods), and training.max_grad_norm (static GRPO).

Locks

config_loader.py, train_sft.py and shared/training_utils.py are offline SFT worker closure members. The closure manifest and Modal runtime lock hashes were refreshed with the checked-in scripts. The member list is unchanged, and all four lock checks report CURRENT.

Behaviour changes for host projects

  • Host GRPO configs copied from the old template that still contain rewards.items[].params or rewards.custom.module will now be refused. Delete those keys.
  • warmup_ratio and weight_decay are no longer in the experiment loop's default search space. They have no trainer flag, and were never actually applied.

Test plan

  • New strictness, GRPO-schema and argv-contract tests pass.
  • Existing KTO/DPO/SFT config, HF Jobs, RTX, local-run and closure tests pass. The remaining failures also occur on the base commit, and come from torch not being installed in the sandbox.
  • Full CI run on the top of the stack (PR 5): core 11,481 passed, 6 failed. All 6 are sandbox-environment artifacts. Both Modal inference shards pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT


Generated by Claude Code

claude added 3 commits October 2, 2026 10:49
Config typos and unsupported settings now fail loudly instead of silently
changing a training run.

Unknown keys:
- shared/training_utils.py gains one shared dict_to_dataclass plus
  reject_unknown_config_keys / find_unknown_config_keys. Every undeclared
  key at any nesting level is listed as a dotted path (e.g.
  training.lerning_rate, rewards.items[0].wieght) with a difflib
  "did you mean" hint, capped at 20 lines.
- SFT/KTO/DPO config loaders drop their three private dict_to_dataclass
  copies and validate the whole YAML tree against the Config dataclasses.
  The protected HF smoke recipe envelope keys are admitted only on the
  protected-smoke path (PROTECTED_RECIPE_ENVELOPE_KEYS).
- Tier presets refuse keys their tier_config_map does not route.
- GRPO and env-GRPO declare GRPO_CONFIG_SCHEMA / ENV_GRPO_CONFIG_SCHEMA
  next to their loaders; a drift test checks them against the keys the
  code reads.
- Modal recipe sections now name unsupported fields.

Unsupported trainer arguments:
- build_trainer_config replaces the GRPOConfig signature filters. Any
  argument set by YAML, extra_args, a method switch (use_gspo ->
  importance_sampling_level) or a CLI override raises with the arg, its
  config path and the installed trl version. Only env-GRPO's internal
  max_prompt_length default (removed in trl 0.28) is named as
  version-dependent and may be omitted.

Dead or unwired settings:
- Wire SFT group_by_length into SFTConfig, apply wandb.project and
  wandb.entity via WANDB_PROJECT/WANDB_ENTITY, wire GRPO max_grad_norm
  and wandb.run_name.
- Remove dead keys from GRPO YAMLs (schema.tool_schema_path, per-item
  params, custom.module) and from env_config.yaml (model.max_seq_length,
  model.chat_template, training.max_prompt_length, env_training.backend,
  runtime.cloud_base_image).

Refresh the offline SFT worker closure and Modal runtime lock hashes for
the edited closure members; update troubleshooting docs and the
fine-tuning skill.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
- train_env_grpo accepts --save-steps/--save-total-limit (written into
  training.save_steps/save_total_limit, which GRPOConfig already reads),
  so HF env-GRPO launches no longer pass flags the trainer rejects.
- Remove the dead train_env_grpo --max-seq-length flag. hf_jobs_backend no
  longer derives max_seq_length for grpo (it fell back to
  max_prompt_length), and the HF command builder refuses an explicit
  max_seq_length override for grpo instead of forwarding it.
- apply_tier_preset coerces tier values through the same
  coerce_config_value path dict_to_dataclass uses, so `5e-4`-style rates
  (loaded by YAML as strings) become floats.
- Delete KTO/DPO dataset.chat_template: declared and set in config.yaml
  but read by nothing. Every other SFT/KTO/DPO field, including the KTO
  use_kto_s / two-stage LR keys, is read.
- Refresh the offline SFT worker closure and Modal runtime lock hashes for
  the shared/training_utils.py change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
- KTO/DPO accept --save-steps / --save-total-limit (applied to
  training.save_steps / save_total_limit), which the HF Jobs builder
  already passes from their configs.
- Env-GRPO accepts --seed and a top-level seed key, passed to
  GRPOConfig.seed; the HF builder passes --seed for seed overrides.
- HF Jobs training refuses methods whose trainer lacks its run/artifact
  flags (embedding, ace_step) instead of launching a job argparse rejects.
- local-run refuses model.revision for kto/dpo (only train_sft pins a Hub
  revision). It already refuses grpo without an explicit run.command, so
  no hyperparameter flags reach train_grpo.py.
- The experiment loop no longer passes its flat override YAML as --config
  (train_sft would exec it as Python and train_kto rejects --config).
  Hyperparameters map to trainer CLI flags via TRAINER_OVERRIDE_FLAGS;
  keys without a flag (warmup_ratio, weight_decay) are refused by
  validate() and at run time, and a missing base_config_path raises.
  Drop those two keys from the default search space.
- KTO's parser moves into build_arg_parser() so it can be loaded without
  the ML stack.

tests/contract/test_trainer_argv_contract.py builds a representative
command with every optional setting for each launcher and method (HF
Jobs, RunPod, local-run including runtime profiles, RTX, Mac, flywheel
orchestrator and experiment loop, checked-in recipe run steps, runtime_v1,
the packaged SFT and protected-smoke invocations and the offline worker
flag allowlist) and parses it with the trainer's real argparse, loaded
from source.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
@ProfSynapse
ProfSynapse merged commit 1ee8b14 into feat/submodule-cloud-api-v1 Oct 2, 2026
4 checks passed
@ProfSynapse
ProfSynapse deleted the claude/config-strictness branch October 6, 2026 19:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants