Register and share your invite link to earn from video plays and referrals.

Larry Dial
@classiclarryd
Technical Staff at Open Athena, working on Marin
Joined May 2024
50 Following    2.3K Followers
New NanoGPT Speedrun WR at 68.0 (-5.8s) from @theonlyglitch_ , with a decrease of QK dim from 128 to 96, a kernel fusion of [QK Norm, RoPE, KeyOffset, paired head layout] into Triton, and moving MLP bwk, QKV fwd/bwk to FP8. This is a heavily involved PR with over 1k lines of triton, 4 different FP8 scaling protocols, and fancy register aware epilogue placement. Two takeaways: 1) From a model perspective, if QK dims smaller than 128 works better at nano scale, then perhaps larger than 128 works better at Hero scale. 2) This Nvidia Engr is super legit.
Show more