็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

Zhicheng Cai
@AiolusZero
PhD @ THU-AIR & Intern @ Bytedance Seed
ๅ‚ๅŠ  July 2026
12 ใƒ•ใ‚ฉใƒญใƒผไธญ    67 ใƒ•ใ‚กใƒณ
Proximal Policy Optimization (PPO) has been the foundational standard for LLM post-training, yet it quietly suffers from a fatal flaw in long-horizon reasoning: exploration collapse. ๐Ÿง We explore the root cause: PPO suffers from a decade-old Geometric Fallacy! Excited to share our paper published in ICML 2026: "Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization" paper: blog: work with: @hello_gensi @ericguoxy @yaqinzhang @haozhou_ai ๐Ÿšจ PPOโ€™s Geometric Mismatch ๐Ÿ”ป PPO-Clip implicitly adopts a Euclidean metric, resulting in a geometric mismatch with the non-uniform Riemannian policy manifold. ๐Ÿ”ป This mismatch pathologically suffocates trust regions for low-probability tokens while overly expanding high-probability ones, polarizing the policy and triggering exploration collapse! ๐Ÿš€ Introducing RIPO (Riemannian Isometric Policy Optimization) ๐Ÿ”น Method: RIPO dynamically aligns isometric trust regions with local geometry, allowing larger updates for low-probability actions while tightening updates for high-probability ones to achieve a balanced exploration-exploitation trade-off. ๐Ÿ”น Bonus: Geometric isometry mathematically induces statistical homoscedasticity, balancing bias-variance trade-off for fast, stable training! ๐Ÿ“Œ Key Results ๐Ÿ”น 35% Avg Gain over GRPO: RIPO consistently refresh SOTA across 7 competition-level mathematical reasoning benchmarks. ๐Ÿ”น 5x Token Efficiency: RIPO reaches GRPO's 200-step performance in just 40 steps with highly stable gradient norms. ๐Ÿ”น Sustained Exploration: RIPO maintains policy entropy at a healthy level throughout late-stage training, preserving policy diversity. ๐Ÿ”น Unlocks Pass@K Scaling Ceiling: On HMMT25, RIPO outperforms GRPO by a massive 50% of Pass@128, breaking the intrinsic ceiling of base models! ๐ŸŒŸ Summary RIPO is simple, principled, and highly effective. By replacing PPO's heuristic Euclidean clip with a dynamic, geometrically aligned clip, it offers a new foundation for RL scaling in LLM reasoning.
ใ‚‚ใฃใจ่ฆ‹ใ‚‹