가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Giles Thomas
@gpjt
Building LLMs, learning in public / Founded and led @PythonAnywhere / PSF Fellow
가입 June 2007
198 팔로잉 중    1.6K 팬
I'd trained some models on 40 tokens per parameter. The Chinchilla paper says that doing that is suboptimal. I wanted to see if that was correct for my training setup -- and it looks like it was :-)
더 보기