Checkpoint loading: require matching weights before restoring state - #2010
Open
peterdsharpe wants to merge 2 commits into
Open
peterdsharpe wants to merge 2 commits into
peterdsharpe wants to merge 2 commits into
Conversation
Contributor
CODEOWNERS review mapCurrent for commit ⏳ @CharlelieLrt — 1 file(s)
⏳ @negin513 — 1 file(s)
No CODEOWNER
Comment |
Collaborator
Author
|
/ok to test 01757c8 |
peterdsharpe
marked this pull request as ready for review
September 19, 2026 21:42
peterdsharpe
requested review from
CharlelieLrt and
negin513
as code owners
September 19, 2026 21:42
Contributor
|
The PR appears safe to merge; no outstanding correctness or repository-rule violations remain. SummaryThis PR makes checkpoint restoration resolve model weights against the selected training checkpoint’s filename index and validates every required weight file before mutating model or training state.
Reviews (2) · Last reviewed commit: "Validate matching checkpoint weights bef..." |
This comment has been minimized.
This comment has been minimized.
Collaborator
Author
|
/ok to test ec50e67 |
Collaborator
Author
peterdsharpe
added a commit
to peterdsharpe/physicsnemo
that referenced
this pull request
Sep 19, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
When the newest model weights are deleted but an older file survives,
load_checkpoint(epoch=None)can combine those older weights with the newest optimizer/scheduler state and report the newest epoch. With multiple requested models, it can also restore the first model before detecting missing weights for a later model.Closes #2012.
Resolve the training checkpoint once and require every requested model's weights at that same filename index. Check all required files before restoring any model or training state; missing weights raise
FileNotFoundErrornaming the training checkpoint and affected models. Automatically numbered saves work even when the training-state payload has noepochkey. Newer orphan weights are not substituted for the selected training checkpoint's weights.Serial and distributed loading share file selection and validation. In distributed mode, rank 0 resolves file availability, broadcasts the result, and every rank validates it before loading begins. Other ranks do not inspect or open checkpoint files. Fresh runs, explicitly absent epochs, and weights-only exports retain their existing behavior.
Validation
device_idduring process-group initialization; checkpoint code and test assertions ran normally. Multi-GPU NCCL/FSDP validation remains for CI.Checklist
Dependencies
None. Based directly on main.