Intelligence Index v4.2 is an acceleration of our v5 roadmap, which is focused on long-horizon tasks that better resemble real-world problem solving, a greater proportion of private test sets to prevent gaming, and reducing saturation. Expect further updates as we continue our v5 rollout!
We are accelerating our rollout because the Intelligence Index needs to measure what AI is capable of to be useful to developers making decisions.
The key changes are adding AA-Briefcase, our long-horizon agentic knowledge work benchmark with a private test set, and GDP.pdf, which measures long-context reasoning across thousands of realistic documents. Both improve realism and AA-Briefcase increases the proportion of private test sets. We've also removed GPQA Diamond: it provides little signal for frontier models given its saturation, and its multiple-choice format doesn't reflect real tasks in those domains.
Link to our methodology:
顯示更多