UK AISI Evaluation of GPT-6-Astra
- No-CoT math time horizon:
"Astra can solve significantly more difficult
math problems in a single forward pass than past models. UK AISI measured Astra’s time horizon at 30.9 minutes compared to 3.6 minutes for
GPT 5.6 Sol"
- CoT Controllability:
"Astra shows a substantial increase in CoT Controllability over GPT 5.6 Sol, following the constraint on 93% of samples
compared to 48%. As with previous models, CoT controllability diminishes significantly for longer stretches of reasoning."
- CoT Legibility:
"Astra reasons in a compressed style, to a greater degree
than GPT 5.6 Sol or GPT 5.5. It is generally possible to understand Astra’s
raw reasoning, although there is an increased frequency of phrases with
unclear meaning. UK AISI expects some, but not all, of these phrases
would be understandable given appropriate context (e.g., the cyber model
spec classification levels)."
- Reasoning Summary Availability:
"During AISI’s evaluations, reasoning
summaries were not consistently provided by the user API, with up to 80% missing on long simulated cyber trajectories. If this remains the case in deployment settings, this could undermine reasoning-based
monitoring of summarized CoT such as what AISI intends to use during
cyber evaluations."
"Overall, UK AISI found that Astra has capabilities that could enable it to evade monitoring. This is due to a greatly increased ability to reason within a single forward-pass, and ability to control the content of its chain of thought (as compared to GPT 5.6 Sol).
Importantly, however, UK AISI did not directly
test if Astra evades monitors successfully and makes no claims about the overall monitorability of the model."