登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Larry Dial
@classiclarryd
Technical Staff at Open Athena, working on Marin
参加 May 2024
48 フォロー中    2.2K ファン
New NanoGPT Speedrun WR at 76.3 (-2.9s) from @Lisennlp , with a radical boost to the final attention layer that drops step count by 7%, called lightweight Dynamically Composable MHA. A secondary local window (112 tokens) is recomputed on the same q/k. After softmax it mixes the per-head attn maps into one unified map, then applies the shared map to each head's values with a per-head scale. Every mix and scale operation is computed per-token from a MUDD MLP. All in a custom kernel. Takes inspiration from Dynamically Composable MHA
もっと見る