forked from NVIDIA/Megatron-LM
-
Notifications
You must be signed in to change notification settings - Fork 16
Pull requests: radixark/Megatron-LM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Support bounded distributed-optimizer main initialization
#86
opened Aug 18, 2026 by
zianglih
Loading…
2 of 6 tasks
fix: restore PP process group on batched pipeline P2P ops
#83
opened Aug 14, 2026 by
Laz4rz
Loading…
TOP dense parity for Qwen3 on Blackwell: flashinfer attention, batch-invariant kernel seam, non-reentrant recompute
#70
opened Jul 20, 2026 by
adrenaline21
Loading…
Handle
step key correctly in dp_reshardable checkpoint save with --optimizer-cpu-offload
#69
opened Jul 16, 2026 by
artkorenev
Loading…
fix(true-on-policy): match SGLang's row-linear k-tile reduction order
#65
opened Jul 9, 2026 by
zihaow211
Loading…
[optim] run plan/metadata coordination over a gloo group
#62
opened Jul 7, 2026 by
yueming-yuan
Loading…
[optim] bucket checkpoint save to avoid CPU memory spike during ckpt saving
#61
opened Jul 6, 2026 by
yueming-yuan
Loading…
Fix GB300 torch_dist checkpoint save crashes from forked local writers
#57
opened Jun 17, 2026 by
zyzshishui
Loading…
6 tasks
Add TV (total-variation) loss option for the MTP draft head
#56
opened Jun 16, 2026 by
ElliotXinqiWang
Loading…
fix(mtp): rename MTP submodule transformer_layer -> mtp_model_layer
#54
opened Jun 8, 2026 by
Zhichenzzz
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.