Repository navigation
Refuse unknown config keys and trainer flags the target does not accept - #162
Merged
ProfSynapse merged 3 commits intoOct 2, 2026
Merged
Conversation
Config typos and unsupported settings now fail loudly instead of silently changing a training run. Unknown keys: - shared/training_utils.py gains one shared dict_to_dataclass plus reject_unknown_config_keys / find_unknown_config_keys. Every undeclared key at any nesting level is listed as a dotted path (e.g. training.lerning_rate, rewards.items[0].wieght) with a difflib "did you mean" hint, capped at 20 lines. - SFT/KTO/DPO config loaders drop their three private dict_to_dataclass copies and validate the whole YAML tree against the Config dataclasses. The protected HF smoke recipe envelope keys are admitted only on the protected-smoke path (PROTECTED_RECIPE_ENVELOPE_KEYS). - Tier presets refuse keys their tier_config_map does not route. - GRPO and env-GRPO declare GRPO_CONFIG_SCHEMA / ENV_GRPO_CONFIG_SCHEMA next to their loaders; a drift test checks them against the keys the code reads. - Modal recipe sections now name unsupported fields. Unsupported trainer arguments: - build_trainer_config replaces the GRPOConfig signature filters. Any argument set by YAML, extra_args, a method switch (use_gspo -> importance_sampling_level) or a CLI override raises with the arg, its config path and the installed trl version. Only env-GRPO's internal max_prompt_length default (removed in trl 0.28) is named as version-dependent and may be omitted. Dead or unwired settings: - Wire SFT group_by_length into SFTConfig, apply wandb.project and wandb.entity via WANDB_PROJECT/WANDB_ENTITY, wire GRPO max_grad_norm and wandb.run_name. - Remove dead keys from GRPO YAMLs (schema.tool_schema_path, per-item params, custom.module) and from env_config.yaml (model.max_seq_length, model.chat_template, training.max_prompt_length, env_training.backend, runtime.cloud_base_image). Refresh the offline SFT worker closure and Modal runtime lock hashes for the edited closure members; update troubleshooting docs and the fine-tuning skill. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
- train_env_grpo accepts --save-steps/--save-total-limit (written into training.save_steps/save_total_limit, which GRPOConfig already reads), so HF env-GRPO launches no longer pass flags the trainer rejects. - Remove the dead train_env_grpo --max-seq-length flag. hf_jobs_backend no longer derives max_seq_length for grpo (it fell back to max_prompt_length), and the HF command builder refuses an explicit max_seq_length override for grpo instead of forwarding it. - apply_tier_preset coerces tier values through the same coerce_config_value path dict_to_dataclass uses, so `5e-4`-style rates (loaded by YAML as strings) become floats. - Delete KTO/DPO dataset.chat_template: declared and set in config.yaml but read by nothing. Every other SFT/KTO/DPO field, including the KTO use_kto_s / two-stage LR keys, is read. - Refresh the offline SFT worker closure and Modal runtime lock hashes for the shared/training_utils.py change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
- KTO/DPO accept --save-steps / --save-total-limit (applied to training.save_steps / save_total_limit), which the HF Jobs builder already passes from their configs. - Env-GRPO accepts --seed and a top-level seed key, passed to GRPOConfig.seed; the HF builder passes --seed for seed overrides. - HF Jobs training refuses methods whose trainer lacks its run/artifact flags (embedding, ace_step) instead of launching a job argparse rejects. - local-run refuses model.revision for kto/dpo (only train_sft pins a Hub revision). It already refuses grpo without an explicit run.command, so no hyperparameter flags reach train_grpo.py. - The experiment loop no longer passes its flat override YAML as --config (train_sft would exec it as Python and train_kto rejects --config). Hyperparameters map to trainer CLI flags via TRAINER_OVERRIDE_FLAGS; keys without a flag (warmup_ratio, weight_decay) are refused by validate() and at run time, and a missing base_config_path raises. Drop those two keys from the default search space. - KTO's parser moves into build_arg_parser() so it can be loaded without the ML stack. tests/contract/test_trainer_argv_contract.py builds a representative command with every optional setting for each launcher and method (HF Jobs, RunPod, local-run including runtime profiles, RTX, Mac, flywheel orchestrator and experiment loop, checked-in recipe run steps, runtime_v1, the packaged SFT and protected-smoke invocations and the offline worker flag allowlist) and parses it with the trainer's real argparse, loaded from source. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
This was referenced Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Config typos and unsupported settings used to change a training run without any error. They now stop the run.
This is PR 1 of 5 in a stack. Merge in order.
Unknown config keys
shared/training_utils.py.5e-4are converted through one shared conversion function._sectionerrors now name the offending fields.Settings dropped because the installed TRL doesn't accept them
GRPOConfigarguments silently. Any argument from YAML,extra_args,use_gspoor the CLI that the installed TRL doesn't accept now raises an error. It names the argument and the installed trl version.max_prompt_length, is listed explicitly in code.Launcher and trainer flag mismatches (found by a new contract test)
--save-steps/--save-total-limitto KTO and DPO, and--seedto env-GRPO, which none of them accepted. Now wired.--model-revisionto KTO and DPO, which have no revision pin. Local-run now refuses it.--config, whichtrain_sftexecuted as Python. Hyperparameters now map to explicit trainer flags, and unmapped keys are refused.tests/contract/test_trainer_argv_contract.pybuilds every launcher's argv (HF Jobs, RunPod, local-run, RTX, Mac, flywheel, checked-in recipes, runtime_v1, protected smoke) and parses it with the trainer's real argument parser.Dead config removed, or wired up
rewards.items[].paramsandrewards.custom.module, which were never read.env_training.backend; the code readsenv_backend.--max-seq-length.dataset.chat_template.group_by_length,wandb.projectandwandb.entity(all methods), andtraining.max_grad_norm(static GRPO).Locks
config_loader.py,train_sft.pyandshared/training_utils.pyare offline SFT worker closure members. The closure manifest and Modal runtime lock hashes were refreshed with the checked-in scripts. The member list is unchanged, and all four lock checks report CURRENT.Behaviour changes for host projects
rewards.items[].paramsorrewards.custom.modulewill now be refused. Delete those keys.warmup_ratioandweight_decayare no longer in the experiment loop's default search space. They have no trainer flag, and were never actually applied.Test plan
🤖 Generated with Claude Code
https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Generated by Claude Code