🚨GPT-6 Astra Pro & Fable 5.1 are now above the human baseline on SimpleBench
For a little bit of context this benchmark was literally built around the everyday reasoning tasks humans were still better at
This gap has basically disappeared now, we aren’t far from AI outperforming even the smartest humans at (human) reasoning tasks