注册并分享邀请链接,可获得视频播放与邀请奖励。

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
加入 March 2020
63 正在关注    2.3K 粉丝
I'm building one of the first eval suites, specifically designed to benchmark models like Jev. Since models like Jev don't generate text, and can't write code or call tools, we need an entirely new kind of eval design. Right now I'm in kinda my favorite phase, where I'm just experimenting with a bunch of different ideas. Nothing to share yet, still lots of experiments to run. Very excited to explore this new kind of model, and help other people figure out if it belongs in their workflow, and if so, where. More to come 🖖
显示更多