๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Dhruv Batra
@DhruvBatra_
Co-founder & Chief Scientist @yutori_ai. Prev: Senior Director leading FAIR Embodied AI @MetaAI and Professor @GeorgiaTech.
๊ฐ€์ž… January 2016
754 ํŒ”๋กœ์ž‰ ์ค‘    21.4K ํŒฌ
๐—ก๐—ฎ๐˜ƒ๐—ถ๐—ด๐—ฎ๐˜๐—ผ๐—ฟ ๐—ป๐Ÿญ.๐Ÿฑ โ€œ๐˜€๐—ผ๐—น๐˜ƒ๐—ฒ๐—ฑโ€ ๐—ข๐—ป๐—น๐—ถ๐—ป๐—ฒ ๐— ๐—ถ๐—ป๐—ฑ๐Ÿฎ๐—ช๐—ฒ๐—ฏ: ๐Ÿต๐Ÿณ.๐Ÿฏ% ๐˜€๐˜‚๐—ฐ๐—ฐ๐—ฒ๐˜€๐˜€ ๐—ฟ๐—ฎ๐˜๐—ฒ. While some teams self-report, this result is independently evaluated and verified by OSU NLP Group @osunlp and Careerflow Human Data Labs. All benchmarks are transient attempts at measuring progress. Ultimately, what matters is how a model performs when people use it. But thereโ€™s a sentiment online that computer-use models arenโ€™t progressing quickly. Not true. In the last year, performance on Online Mind2Web has gone from ~40% success to basically saturated. So whatโ€™s next? Most computer-use/browser-use benchmarks are GUI-only. Models (including Navigator n1.5) now support hybrid actions โ€” UI interactions (click, type, scroll) and programmatic actions (e.g., execute JS). Ultimately, weโ€™re headed to a world where computer-use models โ€œagentifyโ€ the long-tail of the web.
๋” ๋ณด๊ธฐ