I did not see TPUs beating GPUs on decode any time soon, until
@woosuk_k and the cracked team pulled it off with a single Pallas kernel running all of Kimi K3 at 709 tok/s 🤯
First TPU inference megakernel I know of, and it is open source today! The numbers speak for themselves, and the blog is definitely worth your time 🚀