Peak performance across hard benchmarks:
โข Best or joint-best on 5/8 benchmarks (DeepSWE, Chartography, Toolathon, GDP.pdf, SWEFish)
โข Chartography: 48.3 (outperforming Opus 5 & Fable 5)
โข DeepSWE: 74.3
Achieved without Fable 5, Fable 5.1, or GPT-6-Astra in the agent pool.
Details: ๐ก