🧵(1/6) We ran Opus 5 on our cybersecurity benchmark, here are the results:
TLDR
- It finds more vulnerabilities than other frontier models, slightly above GPT-5.6-sol
- Opus 5 is smarter than previous generations, but also works much more than it is asked to
- This "hyperactivity" symptom allows it to find more vulnerabilities, but at a cost: the results are very noisy
Here are the details of our investigation👇