I couldn't find a public performant MXFP8 GEMM on MI355X, so I just asked
@KimiDevs K3 to write one, and it beats torch._scaled_mm by >2x times 🤯
This is an adaptation of the official Gluon MXFP4 example. FlyDSL adaptation is WIP.
@AnushElangovan you guys should try K3 for kernel engineering if you haven't!