Pearl (
@prlnet) has the fastest FP8 MoE Kernels for Blackwell chips!
Group-Matrix-Multiplication is one of the most optimized operations in AI. As a former researcher at
@nvidia I can say firsthand that pushing the performance frontier of MatMuls is... Hard.
Pearl-GEMM is ~5% faster than prior state-of-art Quack (
@tri_dao) and ~42% faster than Flashinfer on B200s.