註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Larry Dial
@classiclarryd
Technical Staff at Open Athena, working on Marin
加入 May 2024
51 正在關注    2.5K 粉絲
New NanoGPT Speedrun WR at 68.0 (-5.8s) from @theonlyglitch_ , with a decrease of QK dim from 128 to 96, a kernel fusion of [QK Norm, RoPE, KeyOffset, paired head layout] into Triton, and moving MLP bwk, QKV fwd/bwk to FP8. This is a heavily involved PR with over 1k lines of triton, 4 different FP8 scaling protocols, and fancy register aware epilogue placement. Two takeaways: 1) From a model perspective, if QK dims smaller than 128 works better at nano scale, then perhaps larger than 128 works better at Hero scale. 2) This Nvidia Engr is super legit.
顯示更多