Register and share your invite link to earn from video plays and referrals.

Hanqing Zhu
@zhu_hanqing666
RL Science & Coding RL Bigrun @xai | Grok 4.5 prev. @UTAustin @GoogleDeepMind @AIatMeta Theory-driven ML, efficient in practice Views my own
Joined February 2024
691 Following    4K Followers
People keep asking me: what's different about optimization in RL? Seemingly nothing โ€” the pre-training stack just works (Adam, even SGD ๐Ÿ‘€ @saagnikkk). Bringing some answers from my last work (sorry for the delay โ€” been cooking ๐Ÿš€). We introduce ISO: Isospectral Optimization: an RLVR-native optimization stack. Built on one simple observation, spectral inheritance: RLVR can reuse the base model's spectrum and acquire new behavior purely through the singular frames. ๐Ÿงฉ Offline: ISO-Merger โ€” consolidates RL experts into one model with no data, no rollouts, no OPD. Checkpoints only. โš™๏ธ Online: ISO-Optimizer โ€” a drop-in wrapper on AdamW / Muon that matches AdamW's accuracy with ~2.7ร— fewer steps on Qwen3-8B-Base. ๐Ÿ“„ ๐ŸŒ ๐Ÿงต๐Ÿ‘‡
Show more