๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Zhicheng Cai
@AiolusZero
PhD @ THU-AIR & Intern @ Bytedance Seed
๊ฐ€์ž… July 2026
12 ํŒ”๋กœ์ž‰ ์ค‘    67 ํŒฌ
Proximal Policy Optimization (PPO) has been the foundational standard for LLM post-training, yet it quietly suffers from a fatal flaw in long-horizon reasoning: exploration collapse. ๐Ÿง We explore the root cause: PPO suffers from a decade-old Geometric Fallacy! Excited to share our paper published in ICML 2026: "Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization" paper: blog: work with: @hello_gensi @ericguoxy @yaqinzhang @haozhou_ai ๐Ÿšจ PPOโ€™s Geometric Mismatch ๐Ÿ”ป PPO-Clip implicitly adopts a Euclidean metric, resulting in a geometric mismatch with the non-uniform Riemannian policy manifold. ๐Ÿ”ป This mismatch pathologically suffocates trust regions for low-probability tokens while overly expanding high-probability ones, polarizing the policy and triggering exploration collapse! ๐Ÿš€ Introducing RIPO (Riemannian Isometric Policy Optimization) ๐Ÿ”น Method: RIPO dynamically aligns isometric trust regions with local geometry, allowing larger updates for low-probability actions while tightening updates for high-probability ones to achieve a balanced exploration-exploitation trade-off. ๐Ÿ”น Bonus: Geometric isometry mathematically induces statistical homoscedasticity, balancing bias-variance trade-off for fast, stable training! ๐Ÿ“Œ Key Results ๐Ÿ”น 35% Avg Gain over GRPO: RIPO consistently refresh SOTA across 7 competition-level mathematical reasoning benchmarks. ๐Ÿ”น 5x Token Efficiency: RIPO reaches GRPO's 200-step performance in just 40 steps with highly stable gradient norms. ๐Ÿ”น Sustained Exploration: RIPO maintains policy entropy at a healthy level throughout late-stage training, preserving policy diversity. ๐Ÿ”น Unlocks Pass@K Scaling Ceiling: On HMMT25, RIPO outperforms GRPO by a massive 50% of Pass@128, breaking the intrinsic ceiling of base models! ๐ŸŒŸ Summary RIPO is simple, principled, and highly effective. By replacing PPO's heuristic Euclidean clip with a dynamic, geometrically aligned clip, it offers a new foundation for RL scaling in LLM reasoning.
๋” ๋ณด๊ธฐ