Skip to content

Checkpoint loading: require matching weights before restoring state - #2010

Open
peterdsharpe wants to merge 2 commits into
NVIDIA:mainfrom
peterdsharpe:pr/checkpoint-missing-model-guard
Open

peterdsharpe wants to merge 2 commits into
NVIDIA:mainfrom
peterdsharpe:pr/checkpoint-missing-model-guard

Conversation

@peterdsharpe

@peterdsharpe peterdsharpe commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

Description

When the newest model weights are deleted but an older file survives, load_checkpoint(epoch=None) can combine those older weights with the newest optimizer/scheduler state and report the newest epoch. With multiple requested models, it can also restore the first model before detecting missing weights for a later model.

Closes #2012.

Resolve the training checkpoint once and require every requested model's weights at that same filename index. Check all required files before restoring any model or training state; missing weights raise FileNotFoundError naming the training checkpoint and affected models. Automatically numbered saves work even when the training-state payload has no epoch key. Newer orphan weights are not substituted for the selected training checkpoint's weights.

Serial and distributed loading share file selection and validation. In distributed mode, rank 0 resolves file availability, broadcasts the result, and every rank validates it before loading begins. Other ranks do not inspect or open checkpoint files. Fresh runs, explicitly absent epochs, and weights-only exports retain their existing behavior.

Validation

  • Single-process checkpoint suite: 33 passed, covering PyTorch and PhysicsNeMo models, deleted/renamed files, older and newer mismatched weights, automatic numbering, multi-model failure without partial restoration, and CPU/CUDA loading. Eight existing tracking/cloud integration cases skipped because optional dependencies are unavailable.
  • Two-process CPU/Gloo DTensor tests: 16 passed, including matching model/optimizer/scheduler state, missing weights, automatic numbering, multi-model failure, weights-only exports, and empty directories. Nonzero-rank file discovery and reads are explicitly forbidden. The local launcher omitted an invalid CPU device_id during process-group initialization; checkpoint code and test assertions ran normally. Multi-GPU NCCL/FSDP validation remains for CI.
  • Twelve new serial regression cases failed on the previous PR head before the fixes. The mixed-epoch reproduction also fails on main. Ruff and all changed-file pre-commit hooks passed.

Checklist

  • I am familiar with the contributing guidelines.
  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.
  • The changelog is up to date with these changes.
  • An issue is linked to this pull request.

Dependencies

None. Based directly on main.

@copy-pr-bot

copy-pr-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions

Copy link
Copy Markdown
Contributor

CODEOWNERS review map

Current for commit 01757c816713. An approval covers every file listed for that owner; one owner is sufficient for shared files.

⏳ @CharlelieLrt — 1 file(s)
  • physicsnemo/utils/checkpoint.py
⏳ @negin513 — 1 file(s)
  • physicsnemo/utils/checkpoint.py

No CODEOWNER

  • CHANGELOG.md
  • test/utils/test_checkpoint.py
  • test/utils/test_checkpoint_distributed.py

Comment /codeowners-info to refresh.

@peterdsharpe

Copy link
Copy Markdown
Collaborator Author

/ok to test 01757c8

@peterdsharpe peterdsharpe added the ci:multi-gpu Run this PR on multiGPU ci label Sep 19, 2026
@peterdsharpe
peterdsharpe marked this pull request as ready for review September 19, 2026 21:42
@greptile-apps

greptile-apps Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Retrigger

The PR appears safe to merge; no outstanding correctness or repository-rule violations remain.

Summary

This PR makes checkpoint restoration resolve model weights against the selected training checkpoint’s filename index and validates every required weight file before mutating model or training state.

  • Prevents combining stale model weights with newer optimizer and scheduler state.
  • Prevents partial restoration when one of several requested models is missing.
  • Shares rank-0 file resolution across distributed workers.
  • Adds serial and distributed regression coverage for missing, mismatched, automatically numbered, and weights-only checkpoints.

Reviews (2) · Last reviewed commit: "Validate matching checkpoint weights bef..."

Comment thread physicsnemo/utils/checkpoint.py Outdated
@greptile-apps

This comment has been minimized.

@peterdsharpe peterdsharpe changed the title Checkpoint loading: reject missing model weights when training state exists Checkpoint loading: require matching weights before restoring state Sep 19, 2026
@peterdsharpe

Copy link
Copy Markdown
Collaborator Author

/ok to test ec50e67

@peterdsharpe

Copy link
Copy Markdown
Collaborator Author

@greptileai

peterdsharpe added a commit to peterdsharpe/physicsnemo that referenced this pull request Sep 19, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci:multi-gpu Run this PR on multiGPU ci

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Checkpoint loading can mix model weights and training state from different epochs

1 participant