RLVR has become the recipe for agentic post-training. But for Computer-Use Agents, the bottleneck is not the algorithm, it is the data. ๐
๐ We introduce CUA-Gym: a scalable, lightweight synthesis engine that turns arbitrary task queries into verifiable RLVR data for computer-use agents. The largest open CUA RLVR dataset to date:
๐ฏ 32,122 verifiable RLVR tasks with programmatic setup scripts + rewards
๐ 110 environments: 16 desktop apps + 94 synthesized mock web apps
๐ Qwen3.5-based CUA models trained with GSPO reach 72.6% on OSWorld-Verified and 56.6% on WebArena
๐ Paper:
๐ Homepage:
๐ค Dataset:
๐ป Codebase:
๐งฉ Environments:
๐งต[1/6]