Skip to content

Add doctor sft-mask and fix SFT loss-mask labels - #165

Merged
ProfSynapse merged 3 commits into
feat/submodule-cloud-api-v1from
claude/sft-mask-doctor-port
Oct 2, 2026
Merged

ProfSynapse merged 3 commits into
feat/submodule-cloud-api-v1from
claude/sft-mask-doctor-port

Conversation

@ProfSynapse

Copy link
Copy Markdown
Owner

Summary

PR 4 of 5. Stacked on #164.

This changes training labels. SFT runs before this PR aren't directly comparable with runs after it. Lineage now records preprocessing contract_version: 2.

python tuner.py doctor sft-mask

Loads only the tokenizer, no model and no GPU. It runs sampled rows through the trainer's real per-row path (materialize_sft_row → shared.sft_preprocessing.materialize_sft_example) and the dataset-contract checks.

Hard failures:

  • rows with zero trained tokens;
  • rows where prompt masking stopped before the end of the prompt render;
  • a trained span that doesn't end on the template's end-of-turn token;
  • a doubled BOS at the start of a sequence;
  • dataset-contract violations;
  • a dropped-row fraction above the threshold.

Warnings and info: truncation rate, p50/p95/max token lengths, and multi-turn rows whose earlier assistant turns are untrained.

It also prints a token-by-token preview of which tokens are masked and which are trained. Output is available as --json. Defaults are in Trainers/sft/configs/mask_doctor.yaml, which is loaded through the strict loader. End-of-turn tokens are found by rendering a probe conversation through the template; no model families are hardcoded.

Label fixes

  1. prompt_completion closes with the template's end-of-turn token instead of eos_token. The raw completion encoding is unchanged. The output is byte-identical when the tokenizer's EOS already is its end-of-turn token, which is the usual case for Qwen instruct models. A template with no detectable end-of-turn token falls back to eos_token, with a log line.
  2. prompt_completion honours completion_only_loss: false and trains the full sequence.
  3. Default render masks tokens after the final end-of-turn, such as the template's trailing newline.
  4. Rows left with zero trained tokens after truncation are dropped on every render path, with a logged count per reason. The run fails if the dropped fraction exceeds the new training.max_dropped_row_fraction (default 0.01).
  5. Rows where prompt masking stopped early are never trained on prompt tokens. They follow the same drop-and-threshold policy, so a template-level mismatch, which hits every row, fails the run.

The trainer and the doctor share one end-of-turn detection function, derive_end_of_turn_tokens.

Locks

SFT closure members changed. Hashes were refreshed with the checked-in scripts, and all four checks report CURRENT.

Reviewer check needed

tests/training/test_chat_template_training_transport.py::test_verified_qwen_template_prompt_completion_boundary_offline was skipped here. It needs SYNAPTIC_TEST_QWEN35_TOKENIZER_ARTIFACT, which wasn't available. It should still pass, because of the byte-identical guarantee above. Please run it once with the artifact before merging.

Test plan

  • test_mask_doctor.py (26) and test_sft_label_fixes.py (17) pass. The label tests use an offline tokenizer whose EOS differs from its end-of-turn token.
  • Smoke run on a checked-in dataset (300 rows) with locally built ChatML tokenizers passes for both the default and prompt_completion renders.
  • Pinned Qwen3.5 transport boundary test, run with the tokenizer artifact.
  • Full CI run at the top of the stack passes (see Refuse unknown config keys and trainer flags the target does not accept #162).

🤖 Generated with Claude Code

https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT


Generated by Claude Code

claude added 3 commits October 2, 2026 11:35
The assistant-only mask is derived by prefix-matching the full render
against the add_generation_prompt render of messages[:-1]. When the two
renders diverge early, masking stops silently and prompt tokens are
trained; rows that do not end with an assistant turn silently fall back
to full-sequence loss; and the prompt_completion path appends eos_token
without checking it is the template's end-of-turn token. Nothing
reported any of this.

- PreparedSFTExample gains descriptive diagnostic fields (untruncated
  length, prompt render length, masked prefix length, prefix-mismatch
  flag, expected token at the divergence, fallback reason). Labels and
  input_ids are unchanged on every path.
- preprocessing.materialize_sft_row is the single per-row hop used by
  prepare_sft_dataset and by the doctor. The dataset-level row contract
  moves into validate_sft_dataset_contract, which prepare_sft_dataset
  and the doctor both call. prepare_sft_dataset can emit per-row mask
  diagnostic columns.
- data_loader.load_raw_sft_dataset is the tokenized path's raw loading
  step, shared with the doctor. Every tokenized SFT preparation now logs
  prefix-mismatch and full-sequence-fallback counts and drops the
  diagnostic columns before training.
- New `python tuner.py doctor sft-mask` (Trainers/sft/src/mask_doctor.py,
  tuner/handlers/sft_mask_doctor_handler.py). Settings resolve like
  train_sft.py: trainer config (--sft-config) < explicit flags. Checks:
  dataset contract, row errors, zero trained tokens, prefix mismatch,
  full-sequence fallback, trained span not ending with the end-of-turn
  token derived from the chat template (eos_token fallback and raw_text
  rows, extras in config), doubled BOS, truncation rate and p50/p95/max
  lengths, untrained earlier assistant turns. Token-by-token previews,
  --json, exit 1 on hard failures and 2 on setup errors. Defaults live
  in Trainers/sft/configs/mask_doctor.yaml.
- Router: `doctor sft-mask` is a DOCTOR_SUBCOMMAND_ROUTES entry dispatched
  lazily by _run_doctor (and covered by iter_route_targets); unknown doctor
  subcommands exit 2 instead of silently running system diagnostics.
- mask_doctor.yaml is loaded through the strict config-key checks
  (MaskDoctorConfig); --sft-config goes through the trainer's strict
  load_config.
- --max-seq-length now defaults to None so an explicit value can be told
  apart from the trainer config; no routed command reads the old default.
- Refresh the offline SFT worker closure manifest and the Modal runtime
  lock hash for the three changed closure members, using the checked-in
  regeneration scripts (hash-only; inventory unchanged).
- Docs: common-tasks, project-reference, fine-tuning skill (synced).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
…op untrainable rows

This intentionally changes SFT training labels (preprocessing contract
version 2); losses from earlier runs are not directly comparable.

- Default render: tokens the chat template emits after the final assistant
  turn's end-of-turn token (for example the newline after the turn
  terminator) are no longer trained. The terminator is derived from the
  template by derive_end_of_turn_tokens, now shared by preprocessing and
  doctor sft-mask (moved from the doctor into shared/sft_preprocessing.py,
  cached per tokenizer). Only the final turn is searched, so a truncated
  reply is never masked by an earlier turn's terminator.
- prompt_completion honours completion_only_loss: with it disabled every
  token is trained instead of the prompt being masked anyway.
- Rows with no supervised tokens left (typically truncation) and rows whose
  assistant-only mask stopped before the end of the prompt render are
  dropped by prepare_sft_dataset on every render path, with a logged count
  per reason. The run fails when dropped rows exceed the new declared
  config key training.max_dropped_row_fraction (default 0.01). Dropping
  (rather than failing on the first row) keeps rare boundary quirks from
  blocking a run, while a template-level mismatch exceeds the threshold
  and fails it; prompt tokens are never trained either way. Authoritative
  prepared-message rows keep their stricter fit-or-fail rule.
- The tokenized data loader realigns grouped-validation group values with
  the rows that remain after drops.
- doctor sft-mask reports the same drop counts: zero_trained_tokens and
  mask_prefix_mismatch become per-row warnings and a new dropped_rows
  check fails above the trainer's threshold. The doctor config no longer
  carries its own end-of-turn probe or extra tokens.
- train_sft records preprocessing contract_version 2 in the lineage.
- Not changed: prompt_completion still closes the completion with
  eos_token_id and encodes it without the template. That contract is
  documented and covered by the artifact-backed boundary qualification for
  a pinned production template, so it is left for a reviewed change.
- Refresh the offline SFT worker closure manifest and the Modal runtime
  lock hash for the changed closure members (inventory unchanged).
- Docs and fine-tuning skill note the label change (synced).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
prompt_completion appended eos_token_id after the raw completion. When a
tokenizer's eos differs from the token its chat template uses to end an
assistant turn, the model learned the wrong stop token. The completion is
now closed with the end-of-turn token from the shared
derive_end_of_turn_tokens (the same derivation the default render and
doctor sft-mask use). The raw completion encoding is unchanged.

- New prompt_completion_terminal_id in shared/sft_preprocessing.py. When
  the template renders no end-of-turn token it falls back to eos_token_id
  and logs the fallback once per tokenizer and template kwargs; with
  neither it fails loudly as before.
- Where eos already is the end-of-turn token, rows are byte-identical to
  the previous construction (test added), so the artifact-backed boundary
  qualification's expected target is unchanged.
- Folded into preprocessing contract_version 2 (nothing has shipped).
- Tests: byte-identity when eos equals the end-of-turn token, the
  differing case (<eos> vs <|im_end|>), the logged eos fallback, and the
  doctor's missing_end_of_turn check (now simulated, since real rows no
  longer produce it).
- Docs and fine-tuning skill describe the new terminal (synced).
- Refresh the offline SFT worker closure manifest and the Modal runtime
  lock hash for the changed closure members (inventory unchanged).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
@ProfSynapse
ProfSynapse changed the base branch from claude/cli-routing-fixes to feat/submodule-cloud-api-v1 October 2, 2026 23:28
@ProfSynapse
ProfSynapse merged commit 72fa4de into feat/submodule-cloud-api-v1 Oct 2, 2026
4 checks passed
@ProfSynapse
ProfSynapse deleted the claude/sft-mask-doctor-port branch October 6, 2026 19:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants