註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Aakash Sabharwal
@aakashsabharwal
加入 September 2009
3.8K 正在關注    1.6K 粉絲
New models this week are much more codebase-aware: they check and break their own tests, and like senior/experienced engineers, they even care about code hygiene. As part of that, I want to shout out some major gains we saw on our SWE Atlas leaderboard @ScaleAILabs. @AnthropicAI is tied for 1st on all three SWE Atlas leaderboards, with the latest Opus and Fable models clustered at the top. But Fable 5.1 feels like a qualitatively different model: more consistent, more efficient, and noticeably more codebase-aware. In Test Writing, it's the first model we've seen that actively tries to mutate code and break its own tests to check for robustness. In Refactoring, it consistently cleans up old leftover code without being asked. On SWE Atlas, no model has cracked 70 percent yet. There's a lot of work these agents still can't do.
顯示更多