註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Benjamin Marie
@bnjmn_marie
Independent AI researcher (LLM, NLP). My blog, The Kaitchup - AI on a Budget:
加入 June 2019
221 正在關注    6.9K 粉絲
One of the most interesting tables in the DeepSeek V4.1 model card. Same model, same benchmark, wildly different scores depending on the agent harness. One more proof that comparing your new model’s score with previously published numbers is close to meaningless unless the harness and config are matched.
顯示更多
0
25
246
18
轉發到社區