Skip to content

Kimi-K3 training support #6872

Description

@yuzhongw-nvidia

This issue tracks Kimi-K3 training support in Megatron Core.

Kimi-K3 combines a hybrid Kimi Delta Attention (KDA) / gated Multi-Head Latent Attention
(MLA) backbone, Attention Residuals (AttnRes), Stable LatentMoE, global-batch Quantile
Balancing (QB), MoonViT-V2, per-head Muon, and native MXFP4 quantization-aware training.

Overall status: KDA and gated MLA are available in dev. AttnRes, Stable LatentMoE, and
global-batch QB are in progress. Full-model composition, MoonViT integration, native MXFP4
training, and the reproducible training recipe remain pending.

Status: ✅ available / merged / resolved · 🚧 in progress / open / draft · 📋 planned / pending validation

Status represents end-to-end capability readiness. A merged supporting PR does not necessarily
mean the entire capability is complete.

Primary scope follows the Kimi-K3 section of the 2026 Q3 Megatron Core MoE roadmap. Lower-priority MoonViT and per-head Muon work, plus native MXFP4 training, are retained here so the tracker covers the intended Kimi-K3 training stack without turning into a complete paper gap list.

Status at a glance

Core functionality

Capability Status Priority Summary
Model architecture 🚧 P0 KDA and gated MLA are merged; AttnRes and Stable LatentMoE are in progress; full composition remains pending
Quantile Balancing 🚧 P0 Baseline QB and the fused TE histogram path are merged; K3 global-batch QB and dense-routing-map integration are open
Per-head Muon 🚧 P1 Per-head Q/gate/K/V orthogonalization for MHA/GQA/MLA is open in #6326
MoonViT and multimodal integration 📋 P1 MoonViT-V2 and the multimodal projector/token-integration path are planned
Training validation and recipe 📋 P0 Full-model correctness, convergence, and a reproducible recipe remain pending

Optimization work

Area Status Priority Current focus
Native MXFP4 training 📋 P1 Add the K3 MXFP4-weight / MXFP8-activation QAT, optimizer, and checkpoint path
KDA kernel fusion 📋 P0.5 Migrate applicable GDN fusion work to KDA and validate the MCore integration
MLA latent context parallelism 📋 P0.5 Implement and validate latent CP for the Kimi-K3 MLA path
CUDA Graphs 📋 P1 Validate KDA execution in composition with the MCore MoE stack

1. Core Functionality

1.1 Model Architecture

Status: 🚧 Core attention modules are available; AttnRes and Stable LatentMoE remain in progress

Component Status Current work
KDA 🚧 First-class hybrid-layer KDA support landed in #6556; lora-proj KDA is in-progress #6877.
Gated MLA ✅ Available in dev Gated MLA and the KDA/MLA hybrid allocation path landed in #6556
Attention Residuals 🚧 Draft Native Block AttnRes with pipeline parallelism and MTP support is proposed in #6840
Stable LatentMoE 🚧 In progress Output normalization (#6449), latent RMSNorm (#6804), and end-to-end SiTU-GLU integration (#6673) are not yet complete
Full model composition 📋 Planned Compose KDA/gated MLA, AttnRes, Stable LatentMoE, MTP, and representative distributed training paths

Stable LatentMoE work

Capability Status Tracking / PRs
Optional post-combine output normalization 🚧 Open #6449
Kimi-K3 latent-MoE RMSNorm 🚧 Open #6804
SiTU-GLU MCore integration 🚧 Open #6673
SiTU-GLU Transformer Engine support ✅ Merged TransformerEngine#3402
SiTU-GLU cuDNN Frontend backend ✅ Merged cudnn-frontend#645
SiTU-GLU backend API-contract fixes 🚧 Open cudnn-frontend#670

1.2 Quantile Balancing

Status: 🚧 Baseline and fused backend available; K3 global-batch integration in progress

Capability Status PR
Baseline Quantile Balancing ✅ Merged #5349
Kimi-K3 global-batch Quantile Balancing 🚧 Open #6637
Dense routing maps for Flex dispatch 🚧 Open #6614
Fused QB router histogram path ✅ Merged TransformerEngine#3395

1.3 Per-head Muon

Status: 📋planned

#6326 adds per-head Q/gate/K/V Muon
orthogonalization, covers MHA/GQA/MLA layouts, and handles tensor parallelism that fragments
query-group blocks. Kimi-K3 optimizer integration and end-to-end recipe validation remain pending.

1.4 MoonViT and Multimodal Integration

Status: 📋planned

  • Add the MoonViT-V2 vision encoder training path.
  • Add the multimodal projector and packed image/video-token integration with the Kimi-K3 language model.

No public MCore implementation PR is currently tracked.

1.5 Training Validation and Recipe

Status: 📋 Planned

  • Validate the complete KDA/gated-MLA, AttnRes, Stable LatentMoE, and QB composition.
  • Exercise representative distributed configurations and checkpoint resume.
  • Publish a reproducible recipe with correctness, memory, performance, and short-convergence evidence.

A concrete model provider and user-facing recipe may live in Megatron Bridge; this tracker records
the Megatron Core capabilities required to run it.

2. Optimization Work

Area Status Plan / dependency
Native MXFP4 training 📋 Planned Support K3's MXFP4 weights / MXFP8 activations from QAT through optimizer update, distributed execution, and checkpointing
KDA kernel fusion 📋 Planned Migrate applicable GDN fusion work to KDA.
MLA latent context parallelism 📋 Draft Implement latent CP and validate correctness, memory, communication, and scaling (#6804).
CUDA Graphs 📋 Planned Validate end-to-end KDA + MCore MoE capture/replay, building on cudnn-frontend#556

References

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions