註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

X Freeze
@XFreeze
加入 July 2024
2K 正在關注    258.7K 粉絲
Grok 4.5 just ranked #1# on VulcanBench, outperforming Claude Fable 5, GPT-5.6 Sol and Kimi K3 across 23 real-world merged PRs in Python, Rust, TypeScript, JavaScript and Go Insane efficiency: • 21/23 tasks solved - 91.3% • Best result across all 11 configurations tested • Just $0.32 per solved task Claude Fable 5 and GPT-5.6 Sol peaked at 20/23. Kimi K3 needed a two-hour extended budget just to reach the same score and cost $1.37 per solved task Grok 4.5 did not just win on accuracy. It won on economics too Grok’s coding efficiency is insane
顯示更多
0
33
166
21
轉發到社區