注册并分享邀请链接,可获得视频播放与邀请奖励。

Larry Dial
@classiclarryd
Technical Staff at Open Athena, working on Marin
加入 May 2024
48 正在关注    2.2K 粉丝
New NanoGPT Speedrun WR at 76.3 (-2.9s) from @Lisennlp , with a radical boost to the final attention layer that drops step count by 7%, called lightweight Dynamically Composable MHA. A secondary local window (112 tokens) is recomputed on the same q/k. After softmax it mixes the per-head attn maps into one unified map, then applies the shared map to each head's values with a per-head scale. Every mix and scale operation is computed per-token from a MUDD MLP. All in a custom kernel. Takes inspiration from Dynamically Composable MHA
显示更多
0
6
301
29
转发到社区