Repository navigation
Add doctor sft-mask and fix SFT loss-mask labels - #165
Merged
ProfSynapse merged 3 commits intoOct 2, 2026
Merged
Conversation
The assistant-only mask is derived by prefix-matching the full render against the add_generation_prompt render of messages[:-1]. When the two renders diverge early, masking stops silently and prompt tokens are trained; rows that do not end with an assistant turn silently fall back to full-sequence loss; and the prompt_completion path appends eos_token without checking it is the template's end-of-turn token. Nothing reported any of this. - PreparedSFTExample gains descriptive diagnostic fields (untruncated length, prompt render length, masked prefix length, prefix-mismatch flag, expected token at the divergence, fallback reason). Labels and input_ids are unchanged on every path. - preprocessing.materialize_sft_row is the single per-row hop used by prepare_sft_dataset and by the doctor. The dataset-level row contract moves into validate_sft_dataset_contract, which prepare_sft_dataset and the doctor both call. prepare_sft_dataset can emit per-row mask diagnostic columns. - data_loader.load_raw_sft_dataset is the tokenized path's raw loading step, shared with the doctor. Every tokenized SFT preparation now logs prefix-mismatch and full-sequence-fallback counts and drops the diagnostic columns before training. - New `python tuner.py doctor sft-mask` (Trainers/sft/src/mask_doctor.py, tuner/handlers/sft_mask_doctor_handler.py). Settings resolve like train_sft.py: trainer config (--sft-config) < explicit flags. Checks: dataset contract, row errors, zero trained tokens, prefix mismatch, full-sequence fallback, trained span not ending with the end-of-turn token derived from the chat template (eos_token fallback and raw_text rows, extras in config), doubled BOS, truncation rate and p50/p95/max lengths, untrained earlier assistant turns. Token-by-token previews, --json, exit 1 on hard failures and 2 on setup errors. Defaults live in Trainers/sft/configs/mask_doctor.yaml. - Router: `doctor sft-mask` is a DOCTOR_SUBCOMMAND_ROUTES entry dispatched lazily by _run_doctor (and covered by iter_route_targets); unknown doctor subcommands exit 2 instead of silently running system diagnostics. - mask_doctor.yaml is loaded through the strict config-key checks (MaskDoctorConfig); --sft-config goes through the trainer's strict load_config. - --max-seq-length now defaults to None so an explicit value can be told apart from the trainer config; no routed command reads the old default. - Refresh the offline SFT worker closure manifest and the Modal runtime lock hash for the three changed closure members, using the checked-in regeneration scripts (hash-only; inventory unchanged). - Docs: common-tasks, project-reference, fine-tuning skill (synced). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
…op untrainable rows This intentionally changes SFT training labels (preprocessing contract version 2); losses from earlier runs are not directly comparable. - Default render: tokens the chat template emits after the final assistant turn's end-of-turn token (for example the newline after the turn terminator) are no longer trained. The terminator is derived from the template by derive_end_of_turn_tokens, now shared by preprocessing and doctor sft-mask (moved from the doctor into shared/sft_preprocessing.py, cached per tokenizer). Only the final turn is searched, so a truncated reply is never masked by an earlier turn's terminator. - prompt_completion honours completion_only_loss: with it disabled every token is trained instead of the prompt being masked anyway. - Rows with no supervised tokens left (typically truncation) and rows whose assistant-only mask stopped before the end of the prompt render are dropped by prepare_sft_dataset on every render path, with a logged count per reason. The run fails when dropped rows exceed the new declared config key training.max_dropped_row_fraction (default 0.01). Dropping (rather than failing on the first row) keeps rare boundary quirks from blocking a run, while a template-level mismatch exceeds the threshold and fails it; prompt tokens are never trained either way. Authoritative prepared-message rows keep their stricter fit-or-fail rule. - The tokenized data loader realigns grouped-validation group values with the rows that remain after drops. - doctor sft-mask reports the same drop counts: zero_trained_tokens and mask_prefix_mismatch become per-row warnings and a new dropped_rows check fails above the trainer's threshold. The doctor config no longer carries its own end-of-turn probe or extra tokens. - train_sft records preprocessing contract_version 2 in the lineage. - Not changed: prompt_completion still closes the completion with eos_token_id and encodes it without the template. That contract is documented and covered by the artifact-backed boundary qualification for a pinned production template, so it is left for a reviewed change. - Refresh the offline SFT worker closure manifest and the Modal runtime lock hash for the changed closure members (inventory unchanged). - Docs and fine-tuning skill note the label change (synced). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
prompt_completion appended eos_token_id after the raw completion. When a tokenizer's eos differs from the token its chat template uses to end an assistant turn, the model learned the wrong stop token. The completion is now closed with the end-of-turn token from the shared derive_end_of_turn_tokens (the same derivation the default render and doctor sft-mask use). The raw completion encoding is unchanged. - New prompt_completion_terminal_id in shared/sft_preprocessing.py. When the template renders no end-of-turn token it falls back to eos_token_id and logs the fallback once per tokenizer and template kwargs; with neither it fails loudly as before. - Where eos already is the end-of-turn token, rows are byte-identical to the previous construction (test added), so the artifact-backed boundary qualification's expected target is unchanged. - Folded into preprocessing contract_version 2 (nothing has shipped). - Tests: byte-identity when eos equals the end-of-turn token, the differing case (<eos> vs <|im_end|>), the logged eos fallback, and the doctor's missing_end_of_turn check (now simulated, since real rows no longer produce it). - Docs and fine-tuning skill describe the new terminal (synced). - Refresh the offline SFT worker closure manifest and the Modal runtime lock hash for the changed closure members (inventory unchanged). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
4 of 5 tasks
ProfSynapse
changed the base branch from
claude/cli-routing-fixes
to
feat/submodule-cloud-api-v1
October 2, 2026 23:28
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
PR 4 of 5. Stacked on #164.
This changes training labels. SFT runs before this PR aren't directly comparable with runs after it. Lineage now records preprocessing
contract_version: 2.python tuner.py doctor sft-maskLoads only the tokenizer, no model and no GPU. It runs sampled rows through the trainer's real per-row path (
materialize_sft_row→shared.sft_preprocessing.materialize_sft_example) and the dataset-contract checks.Hard failures:
Warnings and info: truncation rate, p50/p95/max token lengths, and multi-turn rows whose earlier assistant turns are untrained.
It also prints a token-by-token preview of which tokens are masked and which are trained. Output is available as
--json. Defaults are inTrainers/sft/configs/mask_doctor.yaml, which is loaded through the strict loader. End-of-turn tokens are found by rendering a probe conversation through the template; no model families are hardcoded.Label fixes
prompt_completioncloses with the template's end-of-turn token instead ofeos_token. The raw completion encoding is unchanged. The output is byte-identical when the tokenizer's EOS already is its end-of-turn token, which is the usual case for Qwen instruct models. A template with no detectable end-of-turn token falls back toeos_token, with a log line.prompt_completionhonourscompletion_only_loss: falseand trains the full sequence.training.max_dropped_row_fraction(default 0.01).The trainer and the doctor share one end-of-turn detection function,
derive_end_of_turn_tokens.Locks
SFT closure members changed. Hashes were refreshed with the checked-in scripts, and all four checks report CURRENT.
Reviewer check needed
tests/training/test_chat_template_training_transport.py::test_verified_qwen_template_prompt_completion_boundary_offlinewas skipped here. It needsSYNAPTIC_TEST_QWEN35_TOKENIZER_ARTIFACT, which wasn't available. It should still pass, because of the byte-identical guarantee above. Please run it once with the artifact before merging.Test plan
test_mask_doctor.py(26) andtest_sft_label_fixes.py(17) pass. The label tests use an offline tokenizer whose EOS differs from its end-of-turn token.prompt_completionrenders.🤖 Generated with Claude Code
https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Generated by Claude Code