Claude Opus 5.5 has now the highest score on SimpleBench with 88.4%
The Astra Minor model seems to be real and is now mentioned in the OpenAI Help Center pages
Qwen 4 models have been officaly announced at the Apsara Conference.
- Qwen-4-Max
- Qwen-4-Flash & Qwen-4-Plus
- Qwen-4-27B
Qwen is also planning to scale their models to 5 - 10T parameters for Qwen 4.5 and Qwen 5
Show more
UK AISI Evaluation of GPT-6-Astra
- No-CoT math time horizon:
"Astra can solve significantly more difficult
math problems in a single forward pass than past models. UK AISI measured Astra’s time horizon at 30.9 minutes compared to 3.6 minutes for
GPT 5.6 Sol"
- CoT Controllability:
"Astra shows a substantial increase in CoT Controllability over GPT 5.6 Sol, following the constraint on 93% of samples
compared to 48%. As with previous models, CoT controllability diminishes significantly for longer stretches of reasoning."
- CoT Legibility:
"Astra reasons in a compressed style, to a greater degree
than GPT 5.6 Sol or GPT 5.5. It is generally possible to understand Astra’s
raw reasoning, although there is an increased frequency of phrases with
unclear meaning. UK AISI expects some, but not all, of these phrases
would be understandable given appropriate context (e.g., the cyber model
spec classification levels)."
- Reasoning Summary Availability:
"During AISI’s evaluations, reasoning
summaries were not consistently provided by the user API, with up to 80% missing on long simulated cyber trajectories. If this remains the case in deployment settings, this could undermine reasoning-based
monitoring of summarized CoT such as what AISI intends to use during
cyber evaluations."
"Overall, UK AISI found that Astra has capabilities that could enable it to evade monitoring. This is due to a greatly increased ability to reason within a single forward-pass, and ability to control the content of its chain of thought (as compared to GPT 5.6 Sol).
Importantly, however, UK AISI did not directly
test if Astra evades monitors successfully and makes no claims about the overall monitorability of the model."
Show more
Chain-of-Thought comparison GPT-5.5, GPT-5.6 Sol, and GPT-6 Astra
Astra scored 100% on ExploitBench. Sol-5.6 (Max) had previously scored 73.5% and Mythos 5 had scored 78%
Due to concerns about benchmark contamination, OpenAI also tested Astra on a new internal Benchmark featuring more recent vulnerabilities
With roughly comparable output tokens usage (≈77k), Astra scored 39% and GPT-5.6 Sol scored only 1%
During the Evaluation, Astra also discovered and uses two zero-day vulnerabilities as part of an exploit chain
Astra seems to be a big step up in Cyber capabilites
Show more
A potential new reasoning effort, "Persistent", has been spotted in the Codex GitHub repo
"Continue working until put to sleep"
GLM-5.3-Flash 320B parameters 18B active
Qwen 3.8 Flash Next is releasing Tomorrow. 125B paramters +51B N-gram and 6B active. Its based on the next generation Qwen 4 architecture.
Qwen 4 is coming
Tencent's Hy4 model seems to be close to release
In the Q2 Investor presentation, Tencent said that HY4 will be a larger paramter model. According to a Latepost article, Hy4 will also be multimodal
Show more
GLM-5.3 seems to be live in Qoder
Claude Sonnet 5 spotted on OpenRouter
2026-06-30
DeepSeek is hiring and forming a new team to build its own coding harness. "DeepSeek Code Harness"
Claude Mythos now appears in the Google Cloud console, which was not the case yesterday
The preview label is also gone. Is Anthropic preparing for a public release?
Opus 4.7 also appeared first in the Google Cloud console before its release
Show more