註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

nathan chen
@nathancgy4
learning, entropy-maximizing, opinions
加入 April 2022
730 正在關注    20.7K 粉絲
So many exciting research now shared openly. Enjoy!
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: Tech report: Tech blog:
顯示更多
0
13
174
12
轉發到社區