註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Yuandong Tian
@tydsh
Co-founder of @Recursive_SI. ex-Meta FAIR Director. ex-Google. Reasoning, Optimization and Understanding LLM. Novelist in spare time. PhD in @CMU_Robotics.
加入 December 2009
950 正在關注    46.4K 粉絲
🚨A novel way to do RL in LLM post-training! Inspired by our previous path-not-taken work ( we dig deep into the learning trajectory of RL and find that optimizing singular vectors (i.e., rotation) of weight matrices suffices for good performance in RL. The resulting “isospectral optimization” reaches matched scores with substantially fewer training steps. Great work from @zhu_hanqing666 and the co-authors!
顯示更多
People keep asking me: what's different about optimization in RL? Seemingly nothing — the pre-training stack just works (Adam, even SGD 👀 @saagnikkk). Bringing some answers from my last work (sorry for the delay — been cooking 🚀). We introduce ISO: Isospectral Optimization: an RLVR-native optimization stack. Built on one simple observation, spectral inheritance: RLVR can reuse the base model's spectrum and acquire new behavior purely through the singular frames. 🧩 Offline: ISO-Merger — consolidates RL experts into one model with no data, no rollouts, no OPD. Checkpoints only. ⚙️ Online: ISO-Optimizer — a drop-in wrapper on AdamW / Muon that matches AdamW's accuracy with ~2.7× fewer steps on Qwen3-8B-Base. 📄 🌐 🧵👇
顯示更多
0
7
658
77
轉發到社區