Two years ago, we built OSWorld 1.0 โ the benchmark that became the standard for computer-use agents. Agents now score 83.5% on it. Problem solved?
Not even close.
๐Today we introduce OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks.
What's new:
๐ฏ 108 real-world workflows, each ~1.6 hours โฑ๏ธ for a skilled human
โ๏ธ ~318 tool calls/task vs. ~30 in OSWorld 1.0
๐ Grounded in authentic artifacts & stateful user profiles
โก Captures real phenomena: dynamic environments, streaming interaction, cross-source reasoning, implicit-state inference & more
๐ Best results: Claude Opus 4.8 reaches the highest accuracy at 20.6%, while GPT-5.5 is far more token-efficient but plateaus near 13%. No one is close to solving real computer use.
๐ Homepage:
๐ Paper:
๐ป Code:
๐ค Dataset:
๐งต [1/8]