The new Kimi K3 leverages LatentMoE, a technique developed by
@nvidia in January
What is LatentMoE? In LatentMoE, tokens are projected from the model hidden dimension đ into a smaller latent dimension â for expert routing and computation, which reduces routed parameter loads and all-to-all traffic by a factor of đ/â.
Learn more about it here: