登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
参加 March 2020
63 フォロー中    2.3K ファン
And the benchmark comparison of GPT 5.5 vs. Luna, across all effort levels has begun. Very interested to see the results. The general consensus seems to be that Luna should win on price and accuracy, if that's the case, then there really would be no reason to ever use GPT 5.5 any more. But we'll let the data tell us 🖖
もっと見る