Register and share your invite link to earn from video plays and referrals.

Francesco
@francedot
the CUA wizard ʕ•ᴥ•ʔ @trycua
4K Following    7.7K Followers
Pretty cool 🔥🙌🏽
1/ Today we're introducing Cua-S1-4B-0.2, the first multimodal decision model trained with RLOO on live computer-use tasks, using task-completion rewards. Text and multimodal adapters are available under Apache-2.0:
Show more
Who's going to try to slap this on a Codex/Cowork/etc computer use example first?
1/ You can now try Cua-S1-4B-0.2 in your browser. Thanks to @multimodalart at @huggingface for building the Space! Pick an example and see how the model scores the candidate actions.
Show more
Awesome innovation from the @trycua team! We dont need to burn big model inference just to decide what element of a UI to interact with. Using a local small specialized model to handle execution is the way. Cua-S1 is applying almost the same philosophy as Jev, one layer further down. • screen + candidate actions ↓ S1 ↓ • element • action • abstain
Show more
Jev-class models for computer use are a match made in heaven. Great stuff @francedot is building here. Looking forward to benchmark it.
they went from 0.1 to 0.2 in 1 day we could only wish other labs shipped new models with such large improvements so fast
That’s a bold statement, looking at the three.js recent viral posts it’s crazy how much they are starting to thread in the creative world. Great video @francedot!
Love Cua. They are quick, innovative, and nice! Their tech makes computer-use accessible to everyone.
asked claude opus 5.5 to make a launch video for @trycua with @HyperFrames_ took ~1h of steering, but the result is better than anything we've gotten from motion designers we've previously worked with
Show more
1/ Today we're introducing Cua-S1-4B-0.2, the first multimodal decision model trained with RLOO on live computer-use tasks, using task-completion rewards. Text and multimodal adapters are available under Apache-2.0:
Show more
The best vertical integration arc is starting in the world of computer use!
Computer-use agent experience = capability × robustness × efficiency. As frontier models climb higher on CUA benchmarks, efficiency becomes increasingly attrative, as highlighted by the recent release of #Jev#. With many intermediate steps in everyday computer-use tasks being “local,” fast System-1 models like #CUA-S1# and #Jev# can make agent execution much more efficient. Cool exploration by @francedot and @ddupont808 at @trycua!
Show more
Can't even walk home from Chipotle without witnessing a hit and run and getting recognized 💀
0
92
3.8K
21
Forward to community