forked from NVIDIA/Megatron-LM
-
Notifications
You must be signed in to change notification settings - Fork 10
Pull requests: radixark/Megatron-LM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
TOP dense parity for Qwen3 on Blackwell: flashinfer attention, batch-invariant kernel seam, non-reentrant recompute
#70
opened Jul 20, 2026 by
adrenaline21
Loading…
Handle
step key correctly in dp_reshardable checkpoint save with --optimizer-cpu-offload
#69
opened Jul 16, 2026 by
artkorenev
Loading…
fix(true-on-policy): match SGLang's row-linear k-tile reduction order
#65
opened Jul 9, 2026 by
zihaow211
Loading…
feat(optimizer): NVMe streaming of DistributedOptimizer state
#63
opened Jul 8, 2026 by
yueming-yuan
Loading…
[optim] run plan/metadata coordination over a gloo group
#62
opened Jul 7, 2026 by
yueming-yuan
Loading…
[optim] bucket checkpoint save to avoid CPU memory spike during ckpt saving
#61
opened Jul 6, 2026 by
yueming-yuan
Loading…
Fix GB300 torch_dist checkpoint save crashes from forked local writers
#57
opened Jun 17, 2026 by
zyzshishui
Loading…
6 tasks
Add TV (total-variation) loss option for the MTP draft head
#56
opened Jun 16, 2026 by
ElliotXinqiWang
Loading…
fix(mtp): rename MTP submodule transformer_layer -> mtp_model_layer
#54
opened Jun 8, 2026 by
Zhichenzzz
Loading…
[7/15] Complete true-on-policy MoE EP forward orchestration
#40
opened May 26, 2026 by
maocheng23
Loading…
[5/15] Adapt Megatron grouped experts to SGLang layout
#38
opened May 26, 2026 by
maocheng23
Loading…
[3/15] Select SGLang parity paths from transformer config
#36
opened May 26, 2026 by
maocheng23
Loading…
Previous Next
ProTip!
Adding no:label will show everything without a label.