가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Larry Dial
@classiclarryd
Technical Staff at Open Athena, working on Marin
가입 May 2024
50 팔로잉 중    2.3K 팬
New NanoGPT Speedrun WR at 67.6s (-0.4s) from Jan Varho (jvarho on GitHub). Simple idea: mask impossible continuations. EG: "Hyperparameter" tokenizes to ["Hyper", "param", "eter"]. "eter" can never follow "Hyper" in the tokenized dataset since "Hypereter" tokenizes to ["H", "ype", "re", "ter"]. Yet, during multi-token prediction, "Hyper" learns to predict both "param" and "eter". Explicitly masking "eter" helps improve val loss, and during inference prevents a model from generating token pairs its never seen in training.
더 보기