
Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window.
Input
$ 3
per 1M tokens
Output
$ 15
per 1M tokens
Downloads
837,202
Kimi K3 is a 2.8T parameter model built on our Kimi Delta Attention and Attention Residuals, featuring native vision capabilities and a context window of 1 million tokens. It is the world's first open-source 3T-level model, specifically designed for cutting-edge intelligence across long-context encoding, knowledge work, and reasoning.
Novel Architecture:
Kimi K3 is built upon Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and extends MoE sparsity through the Stable LatentMoE framework, which activates 16 out of 896 experts. Compared to Kimi K2, the overall scaling efficiency has improved by approximately 2.5x.
Long-Context Encoding:
Kimi K3 operates with minimal human supervision and is capable of sustained long-duration engineering development, navigating vast codebases, and coordinating endpoint tools—from GPU kernel optimization and compiler development to vision-in-the-loop game development, CAD, and even chip design.
Intelligent Knowledge Work:
Kimi K3 advances end-to-end knowledge work by leveraging its native multimodal architecture to generate in-depth research through interactive visualizations, widgets and dashboards, as well as dynamic design and video editing.
Native Multimodal and Long Context:
Kimi K3 understands text, images, and video within a single model and supports a context window of 1 million tokens.
Open Frontier Weights:
We release the complete Kimi K3 model weights under the Kimi K3 License, making frontier intelligence publicly available for research, deployment, and further innovation.