Register and share your invite link to earn from video plays and referrals.

Zuxin Liu
@LiuZuxin
Post-training frontiers @OpenAI, PhD @CarnegieMellon | Automating myself…
967 Following    3.6K Followers
I was on call for this run and got paged when the first incident happened. It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human. Mixed feelings. One of those moments where capability and risk showed up at the same time.
Show more
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections
Show more
More intelligence per dollar is always a good day :)
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
Show more
Watching one model resolve 100+ long-standing open problems across mathematics in a matter of weeks is the kind of thing that makes you recalibrate your worldview a little. Honestly, I’m still processing this.
Show more
We’re working with an independent advisory group of mathematicians to help OpenAI responsibly share advances in AI and mathematics. The group will advise on how we assess and communicate new mathematical results, uphold academic and professional standards, and build tools that support mathematical research and learning. Through this work, we want mathematicians to be at the center of shaping how AI supports mathematical understanding and how its benefits reach the wider community.
Show more
Another tip, this tip doesn’t apply to Fable because they have a mandatory 30-day retention🤫
Tip for Moonshot and DeepSeek: don't forget to opt out of data sharing with Anthropic in your privacy settings
I spent much of the weekend talking with the team who did this work. Seb--and everyone else--acted with integrity and generosity throughout. Initially we believed the other team had also solved the problem. We wanted to collaborate and do a joint release. When we learned that they had Euler but not Navier-Stokes, we offered to let them go first, to suggest that they should be the ones to get the prize, and optionally for Tristan to be the lead author on a rewrite of the OpenAI proof. We felt it was challenging to offer the same to Levent (an Anthropic employee), who was not willing to talk or coordinate with us anyway. We were open to other solutions. We would have greatly preferred coordination. We did not rush to publish even though the other team wasn't communicating with us. The team threatened us with unfounded accusations of plagarism. Now that we can see their work, the approaches appear to be different. It is also worth noting that our latest model can solve many, many other math problems. It is true that we tried this because there were rumors on the internet last week that Anthropic's models had solved a millennium problem and we were curious if ours could do it too.
Show more
0
1K
11.3K
632
Forward to community
I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
Show more
0
460
5.7K
490
Forward to community
Most ASI-pilled week yet. Two years ago, high-school olympiad math felt impressive. A year ago, it was IMO gold. Today, we might actually need MillenniumBench. What on earth happens next year?
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Show more
Execution is becoming abundant. Human attention is becoming scarce. I’m increasingly bottlenecked by my own bandwidth: keeping up with many parallel agents, deciding what deserves attention, and steering the promising branches. And almost any new idea is cheap enough to just try. As execution gets cheaper, human judgment becomes higher-leverage.
Show more
Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same.
Show more
Crazy how quickly frontier benchmarks get saturated these days. ARC-AGI-1 took 6 years. ARC-AGI-2 took 16 months. ARC-AGI-3 took just 5 months. 🤯 The frontier is moving faster than we can benchmark it.
Show more
arc-agi-3 is now saturated
Wait, resets earn interest now? 😂
We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can. First one will land in ~ 3 hours. There is still time to create your account if you don't have one.
Show more
Hard to imagine having a model this intelligent, this fast, and this cheap just six months ago. The frontier is moving fast.
We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.
Show more
The only hardware upgrade I actually need.
Meet kbd-1.0-codex-micro, built with @work_louder. Map the buttons and joystick to your workflow, and keep your pinned chats in view. Get yours before stock returns 410.
☀️Very proud of the team for training 5.6! A few of my highlights: - Great front-end aesthetics. E.g., I asked Sol to make a new blog for 5.6 with a celestial theme. - Much better CUA - Can work for much longer and is less lazy But it's not all great: - The writing quality has improved, but it's still bad in artifacts - Too many options with models/reasoning/fast mode/Cerebras/multi-agent We'll fix the above. What else should we improve? PS: belated post as there are so many other good models to train.
Show more
some tips on 5.6 sol: 1. remove old slop. try disabling certain community skills and plugins, especially bundles with 20+ skills. think of 5.6 sol like someone who just grew from senior to staff or senior staff level: prescriptive guidance that used to help becomes micromanagement and makes the work worse. you can always add back what clearly helps. 2. turn on codex memory in settings. give codex feedback on what you like and don’t like, and ask it to remember. 3. you probably don’t need ultra. i do 95% of my work on sol high, sometimes use sol extra high, and have only used sol ultra for a handful of sessions. start with high and only move *up* when you’re not happy with the result.
Show more
@AllanZhou17 and i will be answering some questions about sol's coding capabilities drop your questions, spiciest complaints, and feature requests in the reddit AMA thread🌶️
Just feel the AGI (Asian Girl Idol)
guys I just cancelled my Claude plan I don’t know what happened
This is the model that made me feel like a real chunk of my research workflow can be reliably delegated. I no longer write code and launch jobs manually. My work surface is now just Codex + Slack. My proactive Slack bot also suddenly got much better at handling messy context relevant to me, understanding intent, calling Codex to get work done, and preparing drafts before I even see the messages — all with very little attention from me and saved me tons of time🤖 Frontend got a lot better too. Much easier for me to interact with my agents. A lot of work just got faster and easier ⚡️ RSI is coming.
Show more