가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Nathan Lambert
@natolambert
Open model research @ something new. Prev. co-led Olmo at Ai2. Writes @interconnectsai, wrote
가입 December 2014
945 팔로잉 중    102.4K 팬
An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! Called "Never Give Up"
더 보기