注册并分享邀请链接,可获得视频播放与邀请奖励。

nathan chen
@nathancgy4
learning, entropy-maximizing, opinions
加入 April 2022
730 正在关注    20.7K 粉丝
So many exciting research now shared openly. Enjoy!
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: Tech report: Tech blog:
显示更多
0
13
174
12
转发到社区