๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Alexander Yue
@Alezander9
Physics & CS @ Stanford SLAC | Agents @ Browser Use
๊ฐ€์ž… July 2017
59 ํŒ”๋กœ์ž‰ ์ค‘    2.4K ํŒฌ
We talked with hundreds of users, and turned the hardest real user tasks into a browser-use benchmark We invested heavily into atomic, verified, unambiguous rubrics for LLM judges. This is the best browser agent benchmark ever created ๐Ÿงต
๋” ๋ณด๊ธฐ