fun fact: almost all diagrams in the report are made with TikZ written/edited by K3, who has demonstrated remarkable multi-format programmatic rendering capabilities.
btw, as far as i know, k3 is the 2nd best TikZ master after
@yzhang_cs 😉
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights:
Tech report:
Tech blog:
더 보기