注册并分享邀请链接,可获得视频播放与邀请奖励。

Nathan Lambert
@natolambert
Open model research @ something new. Prev. co-led Olmo at Ai2. Writes @interconnectsai, wrote
加入 December 2014
946 正在关注    102.6K 粉丝
An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! Called "Never Give Up"
显示更多
0
36
744
68
转发到社区