็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

Dhruv Batra
@DhruvBatra_
Co-founder & Chief Scientist @yutori_ai. Prev: Senior Director leading FAIR Embodied AI @MetaAI and Professor @GeorgiaTech.
ๅ‚ๅŠ  January 2016
754 ใƒ•ใ‚ฉใƒญใƒผไธญ    21.4K ใƒ•ใ‚กใƒณ
๐—ก๐—ฎ๐˜ƒ๐—ถ๐—ด๐—ฎ๐˜๐—ผ๐—ฟ ๐—ป๐Ÿญ.๐Ÿฑ โ€œ๐˜€๐—ผ๐—น๐˜ƒ๐—ฒ๐—ฑโ€ ๐—ข๐—ป๐—น๐—ถ๐—ป๐—ฒ ๐— ๐—ถ๐—ป๐—ฑ๐Ÿฎ๐—ช๐—ฒ๐—ฏ: ๐Ÿต๐Ÿณ.๐Ÿฏ% ๐˜€๐˜‚๐—ฐ๐—ฐ๐—ฒ๐˜€๐˜€ ๐—ฟ๐—ฎ๐˜๐—ฒ. While some teams self-report, this result is independently evaluated and verified by OSU NLP Group @osunlp and Careerflow Human Data Labs. All benchmarks are transient attempts at measuring progress. Ultimately, what matters is how a model performs when people use it. But thereโ€™s a sentiment online that computer-use models arenโ€™t progressing quickly. Not true. In the last year, performance on Online Mind2Web has gone from ~40% success to basically saturated. So whatโ€™s next? Most computer-use/browser-use benchmarks are GUI-only. Models (including Navigator n1.5) now support hybrid actions โ€” UI interactions (click, type, scroll) and programmatic actions (e.g., execute JS). Ultimately, weโ€™re headed to a world where computer-use models โ€œagentifyโ€ the long-tail of the web.
ใ‚‚ใฃใจ่ฆ‹ใ‚‹