註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
加入 January 2024
670 正在關注    140.5K 粉絲
Intelligence Index v4.2 is an acceleration of our v5 roadmap, which is focused on long-horizon tasks that better resemble real-world problem solving, a greater proportion of private test sets to prevent gaming, and reducing saturation. Expect further updates as we continue our v5 rollout! We are accelerating our rollout because the Intelligence Index needs to measure what AI is capable of to be useful to developers making decisions. The key changes are adding AA-Briefcase, our long-horizon agentic knowledge work benchmark with a private test set, and GDP.pdf, which measures long-context reasoning across thousands of realistic documents. Both improve realism and AA-Briefcase increases the proportion of private test sets. We've also removed GPQA Diamond: it provides little signal for frontier models given its saturation, and its multiple-choice format doesn't reflect real tasks in those domains. Link to our methodology:
顯示更多
0
37
264
10
轉發到社區