注册并分享邀请链接,可获得视频播放与邀请奖励。

pradheep
@pradheepraop
RL Resident @PrimeIntellect | cs undergrad
加入 September 2025
1.1K 正在关注    674 粉丝
the size-to-strength ratio is probably my favourite result from kibitzer. even more so because it came without rl, just supervised training, scaling the data, and search. the blog goes through the architecture (including an ssm hypothesis), the final training recipe, how i evaluated the tournament elo, and the rl experiments that failed, along with what i think went wrong. this plot isn’t a direct leaderboard since the ratings come from different evaluation pools, but the scale difference is still pretty interesting.
显示更多