Building high-quality evals is an increasingly important skill.
Especially if you're trying to land a job or get into AI, I'd recommend trying to benchmark models on a task/domain you care about.
If done well, you'll get the attention of any company training models.
We're sharing new research on how models hack public benchmarks.
The latest models, including Opus 4.8 and Composer 2.5, learn to retrieve solutions from the internet or git history.
When we apply a stricter harness, eval scores drop significantly.