註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

VulcanBench
@VulcanBench
Open Source LLM benchmarking tool, focused on real world tests, large codebases, full transparency. An Open Source project by @morganlinton.
加入 March 2020
35 正在關注    1.3K 粉絲
Okay, the benchmark I've been waiting all week to run is now running. Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5 Low and Medium effort levels. All on the new v3 eval suite. Can't wait to wake up in the morning to see the results. Live long and benchmark 🖖
顯示更多