𝗡𝗮𝘃𝗶𝗴𝗮𝘁𝗼𝗿 𝗻𝟭.𝟱 “𝘀𝗼𝗹𝘃𝗲𝗱” 𝗢𝗻𝗹𝗶𝗻𝗲 𝗠𝗶𝗻𝗱𝟮𝗪𝗲𝗯: 𝟵𝟳.𝟯% 𝘀𝘂𝗰𝗰𝗲𝘀𝘀 𝗿𝗮𝘁𝗲.
While some teams self-report, this result is independently evaluated and verified by OSU NLP Group
@osunlp and Careerflow Human Data Labs.
All benchmarks are transient attempts at measuring progress. Ultimately, what matters is how a model performs when people use it.
But there’s a sentiment online that computer-use models aren’t progressing quickly.
Not true.
In the last year, performance on Online Mind2Web has gone from ~40% success to basically saturated.
So what’s next?
Most computer-use/browser-use benchmarks are GUI-only. Models (including Navigator n1.5) now support hybrid actions — UI interactions (click, type, scroll) and programmatic actions (e.g., execute JS).
Ultimately, we’re headed to a world where computer-use models “agentify” the long-tail of the web.