注册并分享邀请链接,可获得视频播放与邀请奖励。

Aakash Sabharwal
@aakashsabharwal
加入 September 2009
3.8K 正在关注    1.6K 粉丝
New models this week are much more codebase-aware: they check and break their own tests, and like senior/experienced engineers, they even care about code hygiene. As part of that, I want to shout out some major gains we saw on our SWE Atlas leaderboard @ScaleAILabs. @AnthropicAI is tied for 1st on all three SWE Atlas leaderboards, with the latest Opus and Fable models clustered at the top. But Fable 5.1 feels like a qualitatively different model: more consistent, more efficient, and noticeably more codebase-aware. In Test Writing, it's the first model we've seen that actively tries to mutate code and break its own tests to check for robustness. In Refactoring, it consistently cleans up old leftover code without being asked. On SWE Atlas, no model has cracked 70 percent yet. There's a lot of work these agents still can't do.
显示更多