Excited to see TokenSpeed Kernel featured by
@PyTorch.
Portable kernel APIs enable one runtime across multiple GPU architectures without performance regrets, while keeping runtime logic and backend kernels cleanly separated.
Thanks to the PyTorch for highlighting our work.