๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

LotusDecoder
@LotusDecoder
AI - mind - heart
๊ฐ€์ž… December 2023
2.4K ํŒ”๋กœ์ž‰ ์ค‘    7.4K ํŒฌ
ๅ”ๅœฃ๐Ÿฅน
Our Single-rollout Asynchronous Optimization (SAO), is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench.
๋” ๋ณด๊ธฐ