Issue evidence · Run · NVIDIA/Megatron-LM
GitHub issue: https://github.com/NVIDIA/Megatron-LM/issues/8046
Confirmed from source (not run)
Loading a checkpoint into an already allocated distributed optimizer passes the optimizer's own state to load_state_dict, briefly needing about twice the optimizer state in GPU memory.
Read megatron/core/optimizer/distrib_optimizer.py on current main, and ran a plain PyTorch Adam check on CPU. The memory peak itself was not run: it needs Transformer Engine's FusedAdam and a multi-GPU checkpoint.
Code reading plus a partial local check; the out-of-memory peak was not run.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence