月之暗面在“Kimi K3开放日”上正式公布了Kimi K3的模型权重与完整技术报告,并同步开源了支撑其训练的三项关键基础设施(Infra)技术:MoonEP,FlashKDA和AgentEnv。其中FlashKDA此前已开源,MoonEP和AgentEnv则随本次发布正式面向全球开发者开放。
至此,月之暗面将K3的所有技术链路都开放了。
Kimi K3是一个拥有2.8万亿参数的混合专家(MoE)模型,采用896个路由专家、每个Token激活16个专家的稀疏架构。模型具备原生视觉理解能力,并支持100万Token的超长上下文窗口。
相较上一代Kimi K2.5,K3的参数规模提升约3倍。在算力并不宽裕的条件下,月之暗面凭借Kimi Delta Attention(KDA),Attention Residuals以及MoonEP等技术创新,实现了约2.5倍的规模化效率提升。
月之暗面不仅仅开放了模型本身,也将模型权重、训练方法及底层训练系统一并公开。
Hugging Face CEO Clem发文称赞,Kimi K3在30分钟内以超过4000个赞登顶趋势榜,是“迄今为止最快的发布增长速度”;也有人将Kimi K3称为“本地版Fable 5”;Nebius,Fireworks等海外AI基础设施厂商宣布Day0适配;华为昇腾CANN也宣布对模型实现Day0原生支持。
显示更多
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights:
Tech report:
Tech blog:
显示更多