Today we are releasing GPT-5.6-Cyber.
The model is our first large-scale attempt at directly improving capabilities for advanced cybersecurity tasks such as exploit development.
We are finding it to be really quite strong for accelerating defensive work. We are using it across our stack for red-teaming, and our security researchers have used it to find and patch a huge host of 0-day vulnerabilities in open-source software.
Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating "the message board", model misalignment, and more.
I hope it can answer a lot of the questions folks have, and we will release a full detailed postmortem at a later time!
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
To summarize this week:
- we released general purpose computer using agent
- got beaten by a single human in atcoder heuristics competition
- solved 5/6 new IMO problems with natural language proofs
All of those are based on the same single reinforcement learning system
1/N I’m excited to share that our latest @OpenAI experimental reasoning LLM has achieved a longstanding grand challenge in AI: gold medal-level performance on the world’s most prestigious math competition—the International Math Olympiad (IMO).
today we are introducing codex.
it is a software engineering agent that runs in the cloud and does tasks for you, like writing a new feature of fixing a bug.
you can run many tasks in parallel.
Chain-of-thought reasoning provides a natural avenue for improving model safety.
Today we are publishing a paper on how we train the "o" series of models to think carefully through unsafe prompts:
Can we predict emergent capabilities in GPT-N+1🌌 using only GPT-N model checkpoints, which have random performance on the task?
We propose a method for doing exactly this in our paper “Predicting Emergent Capabilities by Finetuning”🧵
🚨 New Jailbreak Bounty Alert
$1,000 for jailbreaking the hidden CoTs from OpenAI's o1-mini and o1-preview!
No bans. Exclusively on the Gray Swan Arena.
🗓Start Time: October 29th, 1 PM ET
🌐Link:
💬Discord: