登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Nick Haber
@nickhaber
Interactively learning AI, cognitive models, learning tools. Assistant Professor at Stanford.
参加 March 2009
321 フォロー中    1K ファン
Great to have this out! Led by @ZiyuX in collaboration with @setlur_amrith @ChaseBlagden @aviral_kumar2 Encourage exploration in RLVR by basing a reward on the privileged information of a reference solution.
もっと見る
🧵(1/9) Sparse RL for reasoning has an exploration problem. It can only reward solutions the model already stumbles into. On hard problems, that means lots of zeros and very little signal. SFT and self-distillation attack this with reference solutions as targets to match. Instead, we use them as reward scaffolds: a dense signal at both the outcome and process level. Introducing ExpRL: RL-based mid-training that improves exploration by scoring the model’s own attempts against the reference via an LLM judge. What we find: • A stronger policy straight out of mid-training (higher pass@1 and pass@k) • Still ahead after downstream sparse-reward RL • Holds across domains – math & STEM • Scales to a larger policy graded by a smaller judge
もっと見る