가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Larry Dial
@classiclarryd
Technical Staff at Open Athena, working on Marin
가입 May 2024
50 팔로잉 중    2.3K 팬
New NanoGPT Speedrun WR at 68.0 (-5.8s) from @theonlyglitch_ , with a decrease of QK dim from 128 to 96, a kernel fusion of [QK Norm, RoPE, KeyOffset, paired head layout] into Triton, and moving MLP bwk, QKV fwd/bwk to FP8. This is a heavily involved PR with over 1k lines of triton, 4 different FP8 scaling protocols, and fancy register aware epilogue placement. Two takeaways: 1) From a model perspective, if QK dims smaller than 128 works better at nano scale, then perhaps larger than 128 works better at Hero scale. 2) This Nvidia Engr is super legit.
더 보기