Our Single-rollout Asynchronous Optimization (SAO), is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks,
such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench.