注册并分享邀请链接,可获得视频播放与邀请奖励。

Benjamin Marie
@bnjmn_marie
Independent AI researcher (LLM, NLP). My blog, The Kaitchup - AI on a Budget:
加入 June 2019
221 正在关注    6.9K 粉丝
The more I think about it, the worse it gets. We’re not just evaluating a harness–model pair. We’re evaluating an adapter–harness–model triplet. The adapter is what you can easily benchmaxx. Take DeepSWE: Pi running through a basic Pi-to-Pier adapter would perform much worse than the same harness and model running through a carefully tuned adapter. Same model, same harness, different integration, potentially very different results. The adapter is rarely published. We really do have a benchmarking problem.
显示更多
One of the most interesting tables in the DeepSeek V4.1 model card. Same model, same benchmark, wildly different scores depending on the agent harness. One more proof that comparing your new model’s score with previously published numbers is close to meaningless unless the harness and config are matched.
显示更多
0
13
106
4
转发到社区