Almost exactly 6 months later I repeated the same experiment - same setup, but with GPT-6 Astra+Fable 5.1 both at Max effort. Result: 5 improvements on TOP of the improvement below (each new one is exponentially harder). Means models are ~10x smarter than 6 months ago.