Register and share your invite link to earn from video plays and referrals.

Steven Dillmann
@StevenDillmann
AI4Science PhD @Stanford @StanfordAILab Terminal-Bench-Science Lead @terminalbench Research Intern @allen_ai Prev. @Cambridge_Uni, @NASAJPL, @imperialcollege
1.9K Following    1.8K Followers
@harborframework changed how agentic evals are built. Now Harbor Adapters port existing benchmarks onto it, and Harbor-Index distills the best across them into one carefully curated cross-benchmark dataset: high quality, high diversity, high difficulty. Huge congrats @LinShi592021 & co!
Show more
Terminal-Bench-Science 0.1 is the #1# featured benchmark on @AnthropicAI’s new Claude Fable release 🚀
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%. 1/n 👇
Show more
0
57
747
137
Forward to community