We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s.
It fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines.