注册并分享邀请链接,可获得视频播放与邀请奖励。

Arena.ai
@arena
Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring →
加入 March 2023
224 正在关注    229.6K 粉丝
DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison, it retains: - 98% of Hy4 preview’s net improvement, at 73% lower cost - 76% of Kimi K3 (Max)’s performance, at 92% lower cost. Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task. Net improvement over Arena baseline | Median cost/task: - Claude Fable 5.1 (Max): +13.90% | $4.54 - GPT 6 Astra (Max): +11.90% | $4.09 - Claude Opus 5 (Max): +11.09% | $3.52 - Claude Opus 5 (High): +10.49% | $2.24 - Claude Fable 5 (High): +9.03% | $2.19 - Claude Opus 4.8 (High): +7.75% | $1.36 - GPT 5.6 Sol (xHigh): +7.40% | $1.09 - Kimi K3 (Max): +6.39% | $0.77 - Hy4 preview: +4.96% | $0.22 - DeepSeek-V4.1-Flash (Max): +4.87% | $0.06 With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena. Congrats again to the @deepseek_ai team on this release!
显示更多
0
54
1.7K
145
转发到社区