Proximal Policy Optimization (PPO) has been the foundational standard for LLM post-training, yet it quietly suffers from a fatal flaw in long-horizon reasoning: exploration collapse.
🧐 We explore the root cause: PPO suffers from a decade-old Geometric Fallacy!
Excited to share our paper published in ICML 2026:
"Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization"
paper:
blog:
work with:
@hello_gensi @ericguoxy @yaqinzhang @haozhou_ai
🚨 PPO’s Geometric Mismatch
🔻 PPO-Clip implicitly adopts a Euclidean metric, resulting in a geometric mismatch with the non-uniform Riemannian policy manifold.
🔻 This mismatch pathologically suffocates trust regions for low-probability tokens while overly expanding high-probability ones, polarizing the policy and triggering exploration collapse!
🚀 Introducing RIPO (Riemannian Isometric Policy Optimization)
🔹 Method: RIPO dynamically aligns isometric trust regions with local geometry, allowing larger updates for low-probability actions while tightening updates for high-probability ones to achieve a balanced exploration-exploitation trade-off.
🔹 Bonus: Geometric isometry mathematically induces statistical homoscedasticity, balancing bias-variance trade-off for fast, stable training!
📌 Key Results
🔹 35% Avg Gain over GRPO:
RIPO consistently refresh SOTA across 7 competition-level mathematical reasoning benchmarks.
🔹 5x Token Efficiency:
RIPO reaches GRPO's 200-step performance in just 40 steps with highly stable gradient norms.
🔹 Sustained Exploration:
RIPO maintains policy entropy at a healthy level throughout late-stage training, preserving policy diversity.
🔹 Unlocks Pass
@K Scaling Ceiling:
On HMMT25, RIPO outperforms GRPO by a massive 50% of Pass
@128, breaking the intrinsic ceiling of base models!
🌟 Summary
RIPO is simple, principled, and highly effective. By replacing PPO's heuristic Euclidean clip with a dynamic, geometrically aligned clip, it offers a new foundation for RL scaling in LLM reasoning.