注册并分享邀请链接,可获得视频播放与邀请奖励。

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
加入 March 2020
63 正在关注    2.3K 粉丝
The days of using a new frontier model, and defaulting to high effort, are officially over.