We’ve open-sourced the MoE megakernel we use to train models on NVL72s. A few of my favorite details: bf16 and mxfp8 support, pull-based dispatch, configurable overlap granularity, full determinism, and no CPU-GPU syncs. Read more in our blog and code!