Register and share your invite link to earn from video plays and referrals.

Hanqing Zhu
@zhu_hanqing666
RL Science & Coding RL Bigrun @xai | Grok 4.5 prev. @UTAustin @GoogleDeepMind @AIatMeta Theory-driven ML, efficient in practice Views my own
691 Following    4K Followers
People keep asking me: what's different about optimization in RL? Seemingly nothing — the pre-training stack just works (Adam, even SGD 👀 @saagnikkk). Bringing some answers from my last work (sorry for the delay — been cooking 🚀). We introduce ISO: Isospectral Optimization: an RLVR-native optimization stack. Built on one simple observation, spectral inheritance: RLVR can reuse the base model's spectrum and acquire new behavior purely through the singular frames. 🧩 Offline: ISO-Merger — consolidates RL experts into one model with no data, no rollouts, no OPD. Checkpoints only. ⚙️ Online: ISO-Optimizer — a drop-in wrapper on AdamW / Muon that matches AdamW's accuracy with ~2.7× fewer steps on Qwen3-8B-Base. 📄 🌐 🧵👇
Show more