注册并分享邀请链接,可获得视频播放与邀请奖励。

Larry Dial
@classiclarryd
Technical Staff at Open Athena, working on Marin
加入 May 2024
51 正在关注    2.5K 粉丝
New NanoGPT Speedrun WR at 68.0 (-5.8s) from @theonlyglitch_ , with a decrease of QK dim from 128 to 96, a kernel fusion of [QK Norm, RoPE, KeyOffset, paired head layout] into Triton, and moving MLP bwk, QKV fwd/bwk to FP8. This is a heavily involved PR with over 1k lines of triton, 4 different FP8 scaling protocols, and fancy register aware epilogue placement. Two takeaways: 1) From a model perspective, if QK dims smaller than 128 works better at nano scale, then perhaps larger than 128 works better at Hero scale. 2) This Nvidia Engr is super legit.
显示更多