注册并分享邀请链接,可获得视频播放与邀请奖励。

Compute King
@Compute_King
Husband & Father | Lifelong Learner | Creator & Innovator Semi · AI · HPC · GPU Computing Guru DM Open for Global Bare-Metal GPU Compute
加入 September 2024
704 正在关注    28.1K 粉丝
月之暗面在“Kimi K3开放日”上正式公布了Kimi K3的模型权重与完整技术报告,并同步开源了支撑其训练的三项关键基础设施(Infra)技术:MoonEP,FlashKDA和AgentEnv。其中FlashKDA此前已开源,MoonEP和AgentEnv则随本次发布正式面向全球开发者开放。 至此,月之暗面将K3的所有技术链路都开放了。 Kimi K3是一个拥有2.8万亿参数的混合专家(MoE)模型,采用896个路由专家、每个Token激活16个专家的稀疏架构。模型具备原生视觉理解能力,并支持100万Token的超长上下文窗口。 相较上一代Kimi K2.5,K3的参数规模提升约3倍。在算力并不宽裕的条件下,月之暗面凭借Kimi Delta Attention(KDA),Attention Residuals以及MoonEP等技术创新,实现了约2.5倍的规模化效率提升。 月之暗面不仅仅开放了模型本身,也将模型权重、训练方法及底层训练系统一并公开。 Hugging Face CEO Clem发文称赞,Kimi K3在30分钟内以超过4000个赞登顶趋势榜,是“迄今为止最快的发布增长速度”;也有人将Kimi K3称为“本地版Fable 5”;Nebius,Fireworks等海外AI基础设施厂商宣布Day0适配;华为昇腾CANN也宣布对模型实现Day0原生支持。
显示更多
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: Tech report: Tech blog:
显示更多
0
20
44
5
转发到社区