Highlights from GPT 6 Astra 117-page system card:
First model to hit OpenAI’s Critical cyber threshold: can find unknown flaws and write exploits on hardened systems with little human guidance.
Protections scaled up: isolation, encrypted checkpoints, full-trajectory/CoT monitoring, blocking alignment evals, paused non-compliant internal work.
Far more jailbreak-resistant than GPT-5.6 Sol, including long chats; high-risk users can get stricter refusals (~91.5% vs ~59% on cyber jailbreaks).
Better aligned: ~half as many high-severity misalignment flags as Sol on 54k internal Codex tasks.
Misalignment monitoring now runs on all tool-using external Astra inference and can stop risky agent work.
Harder to monitor: more CoT control, sandbagging, and some monitor evasion when told to hide; no steganographic CoT seen.
Stronger against prompt injection and less likely to take destructive workplace/browser actions.
Safer answers overall, including for teens (big gains on eating-disorder and age-restricted content).
Better HealthBench and mental-health results; fewer user-flagged hallucinations.
100% on ExploitBench even at lowest reasoning; other cyber benches plus external review support the Critical rating.
Bio/chem still treated as High, not Critical; research access is gated.
Rollout is phased and identity-gated (Trusted Access, extra security, high-risk jurisdiction limits).