註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Violet X.
@ZiyuX
PhD student @Stanford. Working on LLM-based agents
加入 October 2011
371 正在關注    239 粉絲
🧵(1/9) Sparse RL for reasoning has an exploration problem. It can only reward solutions the model already stumbles into. On hard problems, that means lots of zeros and very little signal. SFT and self-distillation attack this with reference solutions as targets to match. Instead, we use them as reward scaffolds: a dense signal at both the outcome and process level. Introducing ExpRL: RL-based mid-training that improves exploration by scoring the model’s own attempts against the reference via an LLM judge. What we find: • A stronger policy straight out of mid-training (higher pass@1 and pass@k) • Still ahead after downstream sparse-reward RL • Holds across domains – math & STEM • Scales to a larger policy graded by a smaller judge
顯示更多
0
3
137
32
轉發到社區