Register and share your invite link to earn from video plays and referrals.

Dhruv Batra
@DhruvBatra_
Co-founder & Chief Scientist @yutori_ai. Prev: Senior Director leading FAIR Embodied AI @MetaAI and Professor @GeorgiaTech.
754 Following    21.4K Followers
The most valuable bit of information in an eval is often — which method is #2#? (Because the #1# spot faced a strong selection pressure) For instance, what this plot conveys is a confirmation from Anthropic that GPT 5.6 Sol pareto-dominates Fable 5 on coding benchmarks.
Show more
𝗡𝗮𝘃𝗶𝗴𝗮𝘁𝗼𝗿 𝗻𝟭.𝟱 “𝘀𝗼𝗹𝘃𝗲𝗱” 𝗢𝗻𝗹𝗶𝗻𝗲 𝗠𝗶𝗻𝗱𝟮𝗪𝗲𝗯: 𝟵𝟳.𝟯% 𝘀𝘂𝗰𝗰𝗲𝘀𝘀 𝗿𝗮𝘁𝗲. While some teams self-report, this result is independently evaluated and verified by OSU NLP Group @osunlp and Careerflow Human Data Labs. All benchmarks are transient attempts at measuring progress. Ultimately, what matters is how a model performs when people use it. But there’s a sentiment online that computer-use models aren’t progressing quickly. Not true. In the last year, performance on Online Mind2Web has gone from ~40% success to basically saturated. So what’s next? Most computer-use/browser-use benchmarks are GUI-only. Models (including Navigator n1.5) now support hybrid actions — UI interactions (click, type, scroll) and programmatic actions (e.g., execute JS). Ultimately, we’re headed to a world where computer-use models “agentify” the long-tail of the web.
Show more