註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Julian Schrittwieser
@Mononofu
Member of Technical Staff at Anthropic prev AlphaGo, AlphaZero, MuZero, AlphaProof, Gemini RL etc at Google DeepMind
加入 August 2007
128 正在關注    32K 粉絲
I’m a big fan of this style of research report, writing up both successful and failed experiments - papers often only present the just-so story of all successful results, making it hard for new researchers to learn how research is actually done!
顯示更多
the size-to-strength ratio is probably my favourite result from kibitzer. even more so because it came without rl, just supervised training, scaling the data, and search. the blog goes through the architecture (including an ssm hypothesis), the final training recipe, how i evaluated the tournament elo, and the rl experiments that failed, along with what i think went wrong. this plot isn’t a direct leaderboard since the ratings come from different evaluation pools, but the scale difference is still pretty interesting.
顯示更多
0
6
165
9
轉發到社區