登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Larry Dial
@classiclarryd
Technical Staff at Open Athena, working on Marin
参加 May 2024
50 フォロー中    2.3K ファン
New NanoGPT Speedrun WR at 67.6s (-0.4s) from Jan Varho (jvarho on GitHub). Simple idea: mask impossible continuations. EG: "Hyperparameter" tokenizes to ["Hyper", "param", "eter"]. "eter" can never follow "Hyper" in the tokenized dataset since "Hypereter" tokenizes to ["H", "ype", "re", "ter"]. Yet, during multi-token prediction, "Hyper" learns to predict both "param" and "eter". Explicitly masking "eter" helps improve val loss, and during inference prevents a model from generating token pairs its never seen in training.
もっと見る