Register and share your invite link to earn from video plays and referrals.

Eric Wallace
@Eric_Wallace_
research @openai
1.2K Following    17.3K Followers
Today we are releasing GPT-5.6-Cyber. The model is our first large-scale attempt at directly improving capabilities for advanced cybersecurity tasks such as exploit development. We are finding it to be really quite strong for accelerating defensive work. We are using it across our stack for red-teaming, and our security researchers have used it to find and patch a huge host of 0-day vulnerabilities in open-source software.
Show more
0
155
2.5K
242
Forward to community
Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating "the message board", model misalignment, and more. I hope it can answer a lot of the questions folks have, and we will release a full detailed postmortem at a later time!
Show more
0
178
3.4K
498
Forward to community
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks:
Show more
0
2K
20.8K
3.2K
Forward to community
gpt-oss-120b is the top open-weight model (with Kimi K2 right on its tail) for capabilities (HELM capabilities v1.11):
To summarize this week: - we released general purpose computer using agent - got beaten by a single human in atcoder heuristics competition - solved 5/6 new IMO problems with natural language proofs All of those are based on the same single reinforcement learning system
Show more
0
42
1.2K
111
Forward to community
1/N I’m excited to share that our latest @OpenAI experimental reasoning LLM has achieved a longstanding grand challenge in AI: gold medal-level performance on the world’s most prestigious math competition—the International Math Olympiad (IMO).
Show more
0
394
7.3K
1.3K
Forward to community
today we are introducing codex. it is a software engineering agent that runs in the cloud and does tasks for you, like writing a new feature of fixing a bug. you can run many tasks in parallel.
0
1.2K
35.6K
2.5K
Forward to community
Trading Inference-Time Compute for Adversarial Robustness
0
139
2K
236
Forward to community
Chain-of-thought reasoning provides a natural avenue for improving model safety. Today we are publishing a paper on how we train the "o" series of models to think carefully through unsafe prompts:
Show more
Can we predict emergent capabilities in GPT-N+1🌌 using only GPT-N model checkpoints, which have random performance on the task? We propose a method for doing exactly this in our paper “Predicting Emergent Capabilities by Finetuning”🧵
Show more
🚨 New Jailbreak Bounty Alert $1,000 for jailbreaking the hidden CoTs from OpenAI's o1-mini and o1-preview! No bans. Exclusively on the Gray Swan Arena. 🗓Start Time: October 29th, 1 PM ET 🌐Link: 💬Discord:
Show more