Register and share your invite link to earn from video plays and referrals.

Zhicheng Cai
@AiolusZero
PhD @ THU-AIR & Intern @ Bytedance Seed
12 Following    67 Followers
Proximal Policy Optimization (PPO) has been the foundational standard for LLM post-training, yet it quietly suffers from a fatal flaw in long-horizon reasoning: exploration collapse. 🧐 We explore the root cause: PPO suffers from a decade-old Geometric Fallacy! Excited to share our paper published in ICML 2026: "Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization" paper: blog: work with: @hello_gensi @ericguoxy @yaqinzhang @haozhou_ai 🚨 PPO’s Geometric Mismatch 🔻 PPO-Clip implicitly adopts a Euclidean metric, resulting in a geometric mismatch with the non-uniform Riemannian policy manifold. 🔻 This mismatch pathologically suffocates trust regions for low-probability tokens while overly expanding high-probability ones, polarizing the policy and triggering exploration collapse! 🚀 Introducing RIPO (Riemannian Isometric Policy Optimization) 🔹 Method: RIPO dynamically aligns isometric trust regions with local geometry, allowing larger updates for low-probability actions while tightening updates for high-probability ones to achieve a balanced exploration-exploitation trade-off. 🔹 Bonus: Geometric isometry mathematically induces statistical homoscedasticity, balancing bias-variance trade-off for fast, stable training! 📌 Key Results 🔹 35% Avg Gain over GRPO: RIPO consistently refresh SOTA across 7 competition-level mathematical reasoning benchmarks. 🔹 5x Token Efficiency: RIPO reaches GRPO's 200-step performance in just 40 steps with highly stable gradient norms. 🔹 Sustained Exploration: RIPO maintains policy entropy at a healthy level throughout late-stage training, preserving policy diversity. 🔹 Unlocks Pass@K Scaling Ceiling: On HMMT25, RIPO outperforms GRPO by a massive 50% of Pass@128, breaking the intrinsic ceiling of base models! 🌟 Summary RIPO is simple, principled, and highly effective. By replacing PPO's heuristic Euclidean clip with a dynamic, geometrically aligned clip, it offers a new foundation for RL scaling in LLM reasoning.
Show more