가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
가입 March 2020
63 팔로잉 중    2.3K 팬
Super interesting stuff, really like the approach the Harbor/TB team is taking here, and very non-trivial.
@harborframework changed how agentic evals are built. Now Harbor Adapters port existing benchmarks onto it, and Harbor-Index distills the best across them into one carefully curated cross-benchmark dataset: high quality, high diversity, high difficulty. Huge congrats @LinShi592021 & co!
더 보기