Astra scored 100% on ExploitBench. Sol-5.6 (Max) had previously scored 73.5% and Mythos 5 had scored 78%
Due to concerns about benchmark contamination, OpenAI also tested Astra on a new internal Benchmark featuring more recent vulnerabilities
With roughly comparable output tokens usage (≈77k), Astra scored 39% and GPT-5.6 Sol scored only 1%
During the Evaluation, Astra also discovered and uses two zero-day vulnerabilities as part of an exploit chain
Astra seems to be a big step up in Cyber capabilites