๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Bowen Wang
@BowenWangNLP
2nd year Ph.D. student @HKUniversity, Prev. @Tsinghua_Uni. Cooking digital agents at @Alibaba_Qwen, Prev @Kimi_moonshot
๊ฐ€์ž… July 2023
481 ํŒ”๋กœ์ž‰ ์ค‘    1.1K ํŒฌ
RLVR has become the recipe for agentic post-training. But for Computer-Use Agents, the bottleneck is not the algorithm, it is the data. ๐ŸŒ ๐Ÿš€ We introduce CUA-Gym: a scalable, lightweight synthesis engine that turns arbitrary task queries into verifiable RLVR data for computer-use agents. The largest open CUA RLVR dataset to date: ๐ŸŽฏ 32,122 verifiable RLVR tasks with programmatic setup scripts + rewards ๐ŸŒ 110 environments: 16 desktop apps + 94 synthesized mock web apps ๐Ÿ† Qwen3.5-based CUA models trained with GSPO reach 72.6% on OSWorld-Verified and 56.6% on WebArena ๐Ÿ“„ Paper: ๐Ÿ  Homepage: ๐Ÿค— Dataset: ๐Ÿ’ป Codebase: ๐Ÿงฉ Environments: ๐Ÿงต[1/6]
๋” ๋ณด๊ธฐ