登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
参加 March 2020
63 フォロー中    2.3K ファン
Cool idea, honored to be included 🖖
picking an llm = three leaderboards, three winners. arena vs academic vs whatever dropped this week. weekend project: UnifyBench ( look at which models look strongest, then drill into the underlying benches for the real detail. 451 models, 109 sources;
もっと見る