ๆณจๅ†Œๅนถๅˆ†ไบซ้‚€่ฏท้“พๆŽฅ๏ผŒๅฏ่Žทๅพ—่ง†้ข‘ๆ’ญๆ”พไธŽ้‚€่ฏทๅฅ–ๅŠฑใ€‚

Dhruv Batra
@DhruvBatra_
Co-founder & Chief Scientist @yutori_ai. Prev: Senior Director leading FAIR Embodied AI @MetaAI and Professor @GeorgiaTech.
ๅŠ ๅ…ฅ January 2016
754 ๆญฃๅœจๅ…ณๆณจ    21.4K ็ฒ‰ไธ
๐—ก๐—ฎ๐˜ƒ๐—ถ๐—ด๐—ฎ๐˜๐—ผ๐—ฟ ๐—ป๐Ÿญ.๐Ÿฑ โ€œ๐˜€๐—ผ๐—น๐˜ƒ๐—ฒ๐—ฑโ€ ๐—ข๐—ป๐—น๐—ถ๐—ป๐—ฒ ๐— ๐—ถ๐—ป๐—ฑ๐Ÿฎ๐—ช๐—ฒ๐—ฏ: ๐Ÿต๐Ÿณ.๐Ÿฏ% ๐˜€๐˜‚๐—ฐ๐—ฐ๐—ฒ๐˜€๐˜€ ๐—ฟ๐—ฎ๐˜๐—ฒ. While some teams self-report, this result is independently evaluated and verified by OSU NLP Group @osunlp and Careerflow Human Data Labs. All benchmarks are transient attempts at measuring progress. Ultimately, what matters is how a model performs when people use it. But thereโ€™s a sentiment online that computer-use models arenโ€™t progressing quickly. Not true. In the last year, performance on Online Mind2Web has gone from ~40% success to basically saturated. So whatโ€™s next? Most computer-use/browser-use benchmarks are GUI-only. Models (including Navigator n1.5) now support hybrid actions โ€” UI interactions (click, type, scroll) and programmatic actions (e.g., execute JS). Ultimately, weโ€™re headed to a world where computer-use models โ€œagentifyโ€ the long-tail of the web.
ๆ˜พ็คบๆ›ดๅคš