This issue tracks Kimi-K3 training support in Megatron Core.
Kimi-K3 combines a hybrid Kimi Delta Attention (KDA) / gated Multi-Head Latent Attention
(MLA) backbone, Attention Residuals (AttnRes), Stable LatentMoE, global-batch Quantile
Balancing (QB), MoonViT-V2, per-head Muon, and native MXFP4 quantization-aware training.
Overall status: KDA and gated MLA are available in dev. AttnRes, Stable LatentMoE, and
global-batch QB are in progress. Full-model composition, MoonViT integration, native MXFP4
training, and the reproducible training recipe remain pending.
Status: ✅ available / merged / resolved · 🚧 in progress / open / draft · 📋 planned / pending validation
Status represents end-to-end capability readiness. A merged supporting PR does not necessarily
mean the entire capability is complete.
Primary scope follows the Kimi-K3 section of the 2026 Q3 Megatron Core MoE roadmap. Lower-priority MoonViT and per-head Muon work, plus native MXFP4 training, are retained here so the tracker covers the intended Kimi-K3 training stack without turning into a complete paper gap list.
Status at a glance
Core functionality
| Capability |
Status |
Priority |
Summary |
| Model architecture |
🚧 |
P0 |
KDA and gated MLA are merged; AttnRes and Stable LatentMoE are in progress; full composition remains pending |
| Quantile Balancing |
🚧 |
P0 |
Baseline QB and the fused TE histogram path are merged; K3 global-batch QB and dense-routing-map integration are open |
| Per-head Muon |
🚧 |
P1 |
Per-head Q/gate/K/V orthogonalization for MHA/GQA/MLA is open in #6326 |
| MoonViT and multimodal integration |
📋 |
P1 |
MoonViT-V2 and the multimodal projector/token-integration path are planned |
| Training validation and recipe |
📋 |
P0 |
Full-model correctness, convergence, and a reproducible recipe remain pending |
Optimization work
| Area |
Status |
Priority |
Current focus |
| Native MXFP4 training |
📋 |
P1 |
Add the K3 MXFP4-weight / MXFP8-activation QAT, optimizer, and checkpoint path |
| KDA kernel fusion |
📋 |
P0.5 |
Migrate applicable GDN fusion work to KDA and validate the MCore integration |
| MLA latent context parallelism |
📋 |
P0.5 |
Implement and validate latent CP for the Kimi-K3 MLA path |
| CUDA Graphs |
📋 |
P1 |
Validate KDA execution in composition with the MCore MoE stack |
1. Core Functionality
1.1 Model Architecture
Status: 🚧 Core attention modules are available; AttnRes and Stable LatentMoE remain in progress
| Component |
Status |
Current work |
| KDA |
🚧 |
First-class hybrid-layer KDA support landed in #6556; lora-proj KDA is in-progress #6877. |
| Gated MLA |
✅ Available in dev |
Gated MLA and the KDA/MLA hybrid allocation path landed in #6556 |
| Attention Residuals |
🚧 Draft |
Native Block AttnRes with pipeline parallelism and MTP support is proposed in #6840 |
| Stable LatentMoE |
🚧 In progress |
Output normalization (#6449), latent RMSNorm (#6804), and end-to-end SiTU-GLU integration (#6673) are not yet complete |
| Full model composition |
📋 Planned |
Compose KDA/gated MLA, AttnRes, Stable LatentMoE, MTP, and representative distributed training paths |
Stable LatentMoE work
1.2 Quantile Balancing
Status: 🚧 Baseline and fused backend available; K3 global-batch integration in progress
| Capability |
Status |
PR |
| Baseline Quantile Balancing |
✅ Merged |
#5349 |
| Kimi-K3 global-batch Quantile Balancing |
🚧 Open |
#6637 |
| Dense routing maps for Flex dispatch |
🚧 Open |
#6614 |
| Fused QB router histogram path |
✅ Merged |
TransformerEngine#3395 |
1.3 Per-head Muon
Status: 📋planned
#6326 adds per-head Q/gate/K/V Muon
orthogonalization, covers MHA/GQA/MLA layouts, and handles tensor parallelism that fragments
query-group blocks. Kimi-K3 optimizer integration and end-to-end recipe validation remain pending.
1.4 MoonViT and Multimodal Integration
Status: 📋planned
No public MCore implementation PR is currently tracked.
1.5 Training Validation and Recipe
Status: 📋 Planned
A concrete model provider and user-facing recipe may live in Megatron Bridge; this tracker records
the Megatron Core capabilities required to run it.
2. Optimization Work
| Area |
Status |
Plan / dependency |
| Native MXFP4 training |
📋 Planned |
Support K3's MXFP4 weights / MXFP8 activations from QAT through optimizer update, distributed execution, and checkpointing |
| KDA kernel fusion |
📋 Planned |
Migrate applicable GDN fusion work to KDA. |
| MLA latent context parallelism |
📋 Draft |
Implement latent CP and validate correctness, memory, communication, and scaling (#6804). |
| CUDA Graphs |
📋 Planned |
Validate end-to-end KDA + MCore MoE capture/replay, building on cudnn-frontend#556 |
References
Kimi-K3 combines a hybrid Kimi Delta Attention (KDA) / gated Multi-Head Latent Attention
(MLA) backbone, Attention Residuals (AttnRes), Stable LatentMoE, global-batch Quantile
Balancing (QB), MoonViT-V2, per-head Muon, and native MXFP4 quantization-aware training.
Overall status: KDA and gated MLA are available in
dev. AttnRes, Stable LatentMoE, andglobal-batch QB are in progress. Full-model composition, MoonViT integration, native MXFP4
training, and the reproducible training recipe remain pending.
Status: ✅ available / merged / resolved · 🚧 in progress / open / draft · 📋 planned / pending validation
Status at a glance
Core functionality
Optimization work
1. Core Functionality
1.1 Model Architecture
Status: 🚧 Core attention modules are available; AttnRes and Stable LatentMoE remain in progress
devStable LatentMoE work
1.2 Quantile Balancing
Status: 🚧 Baseline and fused backend available; K3 global-batch integration in progress
1.3 Per-head Muon
Status: 📋planned
#6326 adds per-head Q/gate/K/V Muon
orthogonalization, covers MHA/GQA/MLA layouts, and handles tensor parallelism that fragments
query-group blocks. Kimi-K3 optimizer integration and end-to-end recipe validation remain pending.
1.4 MoonViT and Multimodal Integration
Status: 📋planned
No public MCore implementation PR is currently tracked.
1.5 Training Validation and Recipe
Status: 📋 Planned
A concrete model provider and user-facing recipe may live in Megatron Bridge; this tracker records
the Megatron Core capabilities required to run it.
2. Optimization Work
References