Kimi K3: Open Frontier Intelligence
Kimi Team's technical report presents Kimi K3, a 2.8-trillion-parameter native multimodal mixture-of-experts model, covering its hybrid attention architecture, training and agentic post-training recipes, systems infrastructure, evaluations, deployment tradeoffs, and open-weight release.
6 sections ยท 16 lessons
Course outline
Frontier Thesis
- Two Scaling Axes
- Reading K3 Numbers
- Evidence Map
Architecture Mechanics
- Hybrid KDA and MLA
- Bounded Delta Attention
- Attention Residuals
- Stable LatentMoE
- Choose the Bottleneck Fix
Pre-Training Recipe
- Curated Multimodal Data
- Scaling Laws and Schedules
- Million-Token Curriculum
- Audit a Long-Context Recipe
Agentic Post-Training
- SFT to RL to MOPD
- White-Box Agent Harnesses
- Verifiable Task Synthesis
- Design an Agent RL Suite
Systems Infrastructure
- KDA Systems Co-Design
- MoE, RL, and Serving State
- Plan a 1M-Context Deployment
Evaluation Judgment
- Benchmark Reading
- Cost, Cases, and Openness
- Make a Deployment Recommendation