Register and share your invite link to earn from video plays and referrals.

XLANG NLP Lab
@XLangNLP
developing embodied AI agents that empower users to use language to interact with digital and physical environments to carry out real-world tasks.
36 Following    1.6K Followers
💻 Meet Qwen-CUA — our native computer-use agent for (almost) everything. Code, APIs, and computer use are three of the most important interfaces for agents. Today’s models are already highly capable with the first two. Qwen-CUA is built to unlock the third: graphical interfaces designed for people. 👀 Native Perception — screenshots only. No DOM, accessibility tree, or other hidden machine-readable state. 🖱️ Native Interaction — keyboard and mouse events across browsers, desktop apps, and professional software. No task-specific APIs. 🧠 Native Intelligence — maintains long-horizon visual context, verifies progress, and learns from large-scale interactive experience with verifiable outcomes. To make native computer use trainable at scale, we built approximately 40K verifiable tasks and rollout infrastructure with nearly 100K vCPUs, supporting tens of thousands of concurrent environments. Across eight computer-use benchmarks spanning everyday desktop use, long-horizon workflows, personalized computing, scientific research, web interaction, macOS, and adversarial robustness, Qwen-CUA demonstrates strong and broadly competitive capabilities. Scaling the same recipe to Qwen-CUA-Max pushes this frontier further. Not just “clicking the screen” — native computer use unlocks software and workflows that previously required a human at the keyboard. Together with code and APIs, it completes the interface stack for more general agents. Joint work by Qwen Team × XLang Lab. 📖 Technical Report: 💻 Code:
Show more
Congrats the Claude team on the big jump on OSWorld! Worth flagging this 70.6% is partial-credit scoring — under strict binary task-level success (OSWorld 1.0-style), Opus 5 is still only ~30%. We're working on improving task quality so task-level binary success becomes more robust and reliable even on hard tasks too.
Show more
On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art:
Exciting to see progress in long-horizon computer use with the GPT-5.6 release. Glad that OSWorld 2.0 is being used to measure end-to-end computer-use workflows in realistic, complex environments. Computer use is moving fast. Looking forward to more results from the community🙌
Show more
OSWorld 2 is taking CUA evaluation to the next level of complexity and realism. Congrats to the team! Glad to contribute to this important effort.
i still remember the discussion of the case of us visa application and i thought that must be mission impossible for ai and just said u guys must be crazy... now it seems that i am the dumb hahah! anyway, time for frontier models to fight again! bon courage!
Show more
108 real-world, long-horizon computer-use workflows. Average rollout: 318 tool calls. Top frontier agent (Claude Opus 4.8 with max thinking + batched tool calls): 20.6% end-to-end completion (54.8% partial progress). Partial progress is real. Reliable end-to-end computer use is not. Proud to be @XLangNLP's research and data partner on OSWorld 2.0. @qi_zhengyang, @vincentsunnchen, and @fredsala contributed from Snorkel.
Show more
From OSWorld 1.0 to 2.0, we went from minutes (~30 steps) to hours (~318), from single apps to real workflows, from high scores (83%) to hard problems (21%). 1+ year, 20+ people, every task rigorously verified. This is what real cua evaluation takes.🙏 👉
Show more
OSWorld 2.0 is here. This is a new batch of ever solid tasks for CUA. Leading us to the next stage of changing the way we let AI free humans’ labour. Congrats and be proud of the team.
After 15+ months of work, I’m thrilled to finally share OSWorld 2.0 🚀 Huge thanks to all the collaborators, advisors, and friends who supported this project from early ideas to release. Hope OSWorld 2.0 helps push the CUA forward. Excited to hear feedback and discuss more!🙌
Show more
Two years ago, we built OSWorld 1.0 — the benchmark that became the standard for computer-use agents. Agents now score 83.5% on it. Problem solved? Not even close. 🚀Today we introduce OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. What's new: 🎯 108 real-world workflows, each ~1.6 hours ⏱️ for a skilled human ⚙️ ~318 tool calls/task vs. ~30 in OSWorld 1.0 🌍 Grounded in authentic artifacts & stateful user profiles ⚡ Captures real phenomena: dynamic environments, streaming interaction, cross-source reasoning, implicit-state inference & more 📊 Best results: Claude Opus 4.8 reaches the highest accuracy at 20.6%, while GPT-5.5 is far more token-efficient but plateaus near 13%. No one is close to solving real computer use. 🏠 Homepage: 📄 Paper: 💻 Code: 🤗 Dataset: 🧵 [1/8]
Show more
0
16
441
109
Forward to community
Current robot policies overfit specific language templates, handling 'pick and place' but freezing on 'drag it to me ' or 'push it closer to me.' They also lack control over execution: which hand, what approach angle, where to grasp, which path to follow. 🤖 FineVLA make robots steerable : changing instruction alters execution; same task, different phrasing, distinct actions — all faithfully done. 🏠 Homepage: 📄 Paper: 💻Codebase: 🧵[1/6]
Show more
RLVR has become the recipe for agentic post-training. But for Computer-Use Agents, the bottleneck is not the algorithm, it is the data. 🐌 🚀 We introduce CUA-Gym: a scalable, lightweight synthesis engine that turns arbitrary task queries into verifiable RLVR data for computer-use agents. The largest open CUA RLVR dataset to date: 🎯 32,122 verifiable RLVR tasks with programmatic setup scripts + rewards 🌐 110 environments: 16 desktop apps + 94 synthesized mock web apps 🏆 Qwen3.5-based CUA models trained with GSPO reach 72.6% on OSWorld-Verified and 56.6% on WebArena 📄 Paper: 🏠 Homepage: 🤗 Dataset: 💻 Codebase: 🧩 Environments: 🧵[1/6]
Show more
Big update for OpenCUA! OpenCUA-72B-preview now ranks #1# on the OSWorld-Verified leaderboard ( It is a pure GUI action, end-to-end computer-use foundation model (Website: Huge thanks to the effort of OpenCUA team and the great support of Kimi Team @Kimi_Moonshot ! Claude 4.5 is extremely strong on OSWorld, but we’re committed to pushing open-source, end-to-end CUA foundation models forward. Over the last month we trained a larger, stronger model: 45.0% average on OSWorld-Verified. It also shows strong GUI grounding ability: 37.3% on UI-Vision @EdwardJian2 @PShravannayak and 60.8% on ScreenSpot-Pro. We’ll keep driving open-source CUA: models will be on HuggingFace very soon, and a paper update is on the way. #OpenSource# #Agents# #OSWorld# #CUA# #ComputerUseAgent#
Show more