註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Giles Thomas
@gpjt
Building LLMs, learning in public / Founded and led @PythonAnywhere / PSF Fellow
加入 June 2007
198 正在關注    1.6K 粉絲
I'd trained some models on 40 tokens per parameter. The Chinchilla paper says that doing that is suboptimal. I wanted to see if that was correct for my training setup -- and it looks like it was :-)
顯示更多