๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Steven Dillmann
@StevenDillmann
AI4Science PhD @Stanford @StanfordAILab Terminal-Bench-Science Lead @terminalbench Research Intern @allen_ai Prev. @Cambridge_Uni, @NASAJPL, @imperialcollege
๊ฐ€์ž… January 2020
1.9K ํŒ”๋กœ์ž‰ ์ค‘    1.8K ํŒฌ
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%. 1/n ๐Ÿ‘‡
๋” ๋ณด๊ธฐ