Register and share your invite link to earn from video plays and referrals.

Cua
@trycua
Close the loop with Cua Driver: /bin/bash -c "$(curl -fsSL "
2.4K Following    20.3K Followers
1/ Today we're announcing the stable Cua Driver release for Omarchy - a new foundation for computer use, built into the OS from the ground up. Over the last month, we worked directly with @dhh, @SpencerGBull and @vaxryy to bring a native synthetic cursor to Omarchy's Hyprland compositor, enabling true multi-cursor computer use at the OS level. Cua Driver is open source and available at
Show more
0
103
1.6K
145
Forward to community
Pretty cool 🔥🙌🏽
1/ Today we're introducing Cua-S1-4B-0.2, the first multimodal decision model trained with RLOO on live computer-use tasks, using task-completion rewards. Text and multimodal adapters are available under Apache-2.0:
Show more
Who's going to try to slap this on a Codex/Cowork/etc computer use example first?
Cua just released Cua-S1-4B-0.2 today. Here's what you need to know. The Cua team, makers of the computer-use driver behind Hermes Agent, has open-sourced this multimodal decision model under Apache-2.0. It's described as the first multimodal decision model trained with RLOO (Reward-Weighted Leave-One-Out) on live computer-use tasks, using task-completion rewards. The model runs on a frozen Qwen3.5-4B base with LoRA adapters. At each step it takes in the screen state, the task goal, and a fixed set of candidate actions, then returns one action, no accessibility tree required. Training happened in two stages: supervised fine-tuning to teach the decision format, then agentic RL in live cua-bench-basic environments where only a fully completed task earns reward. On the GUI-360 benchmark, it scored 92.9% versus 60.1% for comparison. Text and multimodal adapters, training code, and benchmark results are all merged into the Cua GitHub repo and live on Hugging Face. This is version 0.2, up from 0.1 released the day before. Key numbers: - 4B parameters (Qwen3.5-4B base) - 92.9% vs 60.1% on GUI-360 benchmark - Apache-2.0 license - v0.1 to v0.2 in one day The model is small enough to run on-prem, and developers are already testing it on local hardware for business workflows.
Show more
1/ You can now try Cua-S1-4B-0.2 in your browser. Thanks to @multimodalart at @huggingface for building the Space! Pick an example and see how the model scores the candidate actions.
Show more
Awesome innovation from the @trycua team! We dont need to burn big model inference just to decide what element of a UI to interact with. Using a local small specialized model to handle execution is the way. Cua-S1 is applying almost the same philosophy as Jev, one layer further down. • screen + candidate actions ↓ S1 ↓ • element • action • abstain
Show more
Jev-class models for computer use are a match made in heaven. Great stuff @francedot is building here. Looking forward to benchmark it.
@trycua This is a great direction. Not every action needs a giant model thinking through it from scratch. Smaller, purpose built decision models for specific tasks make a ton of sense! 🫡
Cua introduces the Cua-S1-4B-0.2, a multimodal decision model designed for live computer-use tasks. @trycua reached Top 3 on Launch Archive today.
they went from 0.1 to 0.2 in 1 day we could only wish other labs shipped new models with such large improvements so fast
Love the work cua is doing esp with this new release. If you haven't tried computer use models in 6 months worth trying again because they are really good. And if cua does a really good job, then ALL software with various interfaces get abstraced and we can expect massive productivity gains across the economy as a ton of software like EHRs, legacy windows desktop apps, etc become newly accessible to AI.
Show more
Shit is moving so damn fast 😂☠️
4B computer-use model. LoRA adapters on a frozen Qwen3.5-4B, SFT then RL where the only reward is finishing the task. 92.9% vs 60.1% on GUI-360, no accessibility tree. Small enough to run on-prem for real SMB workflows. Pulling it onto the Spark this week
Show more
@trycua Love how fast cua team innovate computer use !
Love Cua. They are quick, innovative, and nice! Their tech makes computer-use accessible to everyone.
1/ Today we're introducing Cua-S1-4B-0.2, the first multimodal decision model trained with RLOO on live computer-use tasks, using task-completion rewards. Text and multimodal adapters are available under Apache-2.0:
Show more
0
89
1.5K
118
Forward to community