๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Laude Institute
@LaudeInstitute
Laude Institute backs computer science researchers turning research into real-world impact. // @LaudeVentures
๊ฐ€์ž… November 2024
426 ํŒ”๋กœ์ž‰ ์ค‘    4.4K ํŒฌ
A living frontier needs a living benchmark. Congratulations to @alexgshaw, @ryan_marten, and the @terminalbench, @frontierbench, and @harborframework teams on this next era, and to our very first Slingshot for continuing to shape how the field measures progress. ๐Ÿ‘
๋” ๋ณด๊ธฐ
Weโ€™re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
๋” ๋ณด๊ธฐ