Register and share your invite link to earn from video plays and referrals.

Hanqing Zhu
@zhu_hanqing666
RL Science & Coding RL Bigrun @xai | Grok 4.5, 4.6 prev. @UTAustin @GoogleDeepMind @AIatMeta Theory-driven ML, efficient in practice Views my own
559 Following    4.4K Followers
FSD is genuinely one of those “wow” moments for me. It has basically given me back my daily commute — and even some of the longer drives I’ve had this week 🫣. Time that used to be spent purely driving can now sometimes be used to catch up on work, deal with an urgent PR when a training issue needs attention, or simply enjoy the trip itself instead of being drained by the drive. That’s what really matters to me about AI: giving humans time and productivity back. Really amazing stuff from @Tesla_AI . Also very grateful to be working with them to make Grok more useful.
Show more
People keep asking me: what's different about optimization in RL? Seemingly nothing — the pre-training stack just works (Adam, even SGD 👀 @saagnikkk). Bringing some answers from my last work (sorry for the delay — been cooking 🚀). We introduce ISO: Isospectral Optimization: an RLVR-native optimization stack. Built on one simple observation, spectral inheritance: RLVR can reuse the base model's spectrum and acquire new behavior purely through the singular frames. 🧩 Offline: ISO-Merger — consolidates RL experts into one model with no data, no rollouts, no OPD. Checkpoints only. ⚙️ Online: ISO-Optimizer — a drop-in wrapper on AdamW / Muon that matches AdamW's accuracy with ~2.7× fewer steps on Qwen3-8B-Base. 📄 🌐 🧵👇
Show more