注册并分享邀请链接,可获得视频播放与邀请奖励。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
加入 January 2024
670 正在关注    140.5K 粉丝
Intelligence Index v4.2 is an acceleration of our v5 roadmap, which is focused on long-horizon tasks that better resemble real-world problem solving, a greater proportion of private test sets to prevent gaming, and reducing saturation. Expect further updates as we continue our v5 rollout! We are accelerating our rollout because the Intelligence Index needs to measure what AI is capable of to be useful to developers making decisions. The key changes are adding AA-Briefcase, our long-horizon agentic knowledge work benchmark with a private test set, and GDP.pdf, which measures long-context reasoning across thousands of realistic documents. Both improve realism and AA-Briefcase increases the proportion of private test sets. We've also removed GPQA Diamond: it provides little signal for frontier models given its saturation, and its multiple-choice format doesn't reflect real tasks in those domains. Link to our methodology:
显示更多
0
37
264
10
转发到社区