登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
参加 March 2020
57 フォロー中    2K ファン
It's official, Astra has saturated some of the most advanced benchmarks out there. Excited to test it on our new eval suite once we have access. Really hoping we can stump it, but also happy for the challenge of iterating and building new evals for the beast of a model Astra sounds like it is. Live long and benchmark 🖖
もっと見る