註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Henry Kiss Ehrenberg
@henryehrenberg
co-founder + engineering @SnorkelAI
加入 February 2013
139 正在關注    393 粉絲
We expect agents to act like senior engineers, but most benchmarks still evaluate them like interns. Excited to introduce Senior SWE-Bench, an open-source and @harborframework-native benchmark that assesses agents as senior engineers on long-horizon tasks with realistically under-specified instructions. We expect agents to build real features going on just a quick Slack message, nothing like the super technical instructions most benchmarks provide. Senior SWE-Bench fixes that. Claude Opus 4.8 is the current leader at 24% high quality solves, but it took 117K tokens on average to get there. Claude Sonnet 5 looked like it was going to swoop in for the top spot, but we found it cheated on 26% of trials.
顯示更多
0
14
239
61
轉發到社區