GPT-6 Astra appears to be a massive jump in opaque reasoning ability: it looks like it can solve hard competition math problems entirely in its head (as in, without verbalized reasoning) while prior AIs could solve basic word problems.
This seems extremely concerning!
That is, if these benchmark results are representative (see the highlighted caveats in the image, I'm particularly worried about contamination).
Related to this, UK AISI found Astra has much worse monitorability.
I'd guess this jump is downstream of architectural changes (with increased serial depth) though a normal large pretrain scale up is a plausible cause. If the next few model generations involve similar jumps (presumably these jumps would be downstream of a transition to full-on opaque reasoning architectures with extreme depth), then chain-of-thought would no longer be a meaningful oversight tool.