There is an extensive and ongoing review related to our agents’ use of internet access during training and evaluation. We’ve been publishing summaries at the link below and will continue to.
We have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations.
We are prioritizing as best as we can based on severity, and adding resources. Hugging Face is still the most severe event we’ve seen. We will be as transparent as we can be subject to things like vulnerabilities in other companies that our agents have found, which will be their call to disclose or not.
Show more
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing.
The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service.
While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties.
Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete.
Show more
argh i meant VOICE not video :(
sorry to dissapoint
Startups are naturally good at this; it is hard to keep a bigger company good at this and i think an underexplored space.
michelle embodies this as much as anyone
openai is so, so lucky to have benefited from everything she has done so far and i think people will be quite pleased to see what she + team have cooking next!
Show more
just crossed four years at openai!
the special thing about this place is the constant capacity for rebirth. for all its faults, there is nowhere quite like it. the team makes high conviction, contrarian bets over and over, and they are mostly right. it's a new company every three months.
i have really enjoyed my part in some of the bets, and i'm excited to share some of the current ones we are cooking up now. onward!
Show more
We want the OpenAI API to feature the best model at every price point and to be the best at every modality (text, code, image, video, etc).
And then we want you all to come up with great ideas and build them and to get to be happy users. The best ideas will come from you all.
Show more
Especially compared by per-task pricing, which is the metric that should matter, I don't think there is anything competitive anywhere in the market.
We want people to be able to use tons of AI; it is important to being able to explore this new renaissance in front of us.
Show more
GPT-6 Sol and Luna are big improvements on intelligence, alignment, work output, coding, computer use, and more over their 5.6-family predecessors.
They are also half the price per token, and even less per task!
Show more
GPT-6 Sol and Luna are great models but also these characters are so cute
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe.
GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale.
We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
Show more
We have been focusing on efficiency and intelligence for all.
Very proud of the team. Only possible when you have incredible models at the top end of the capability that you can then use to make a big difference in everything else.
Show more
People outside the AI labs should have a real say in how this technology develops, and a clear way to judge if it's happening safely.
Standards should help prevent the concentration of power, including by making sure new companies and open-model companies can compete.
They should also help countries and companies compare evidence and learn from failures.
We think the US should lead this effort. Here is our proposal:
Show more
Verified U.S. clinicians can get free GPT-6 Astra (Pro) access today via ChatGPT for Clinicians.
Sign up:
the main thing i was excited about launching this week will be next week instead, but imo worth the wait!
big 🚢 this week
and then for devday
🚢🚢🚢🚢🚢🚢
There are two ways AI progress could go very badly and that we must avoid.
First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities.
Second, we could end up in a world with too much concentration of power. If an extraordinarily powerful AI is used by one person or company to impress their worldview onto everyone else, the results could be extremely dystopian.
Avoiding these two threats requires walking a narrow middle path; for example, one country could gain too much power. Another example is one lab ending up with too much power.
Show more
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot.
We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors).
Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process.
Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases.
We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these.
When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs.
Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring.
Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
Show more
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
Show more
Welcome, Paul.
Grateful you are doing this, and all you have done for AI safety. Excited to work together again.
ocarina of time remake is the best news in a long time.
i will be unavailable november 5 and 6; need to find a case of mountain dew.
This would suck, but we will prioritize great service for customers until we can get back on top of things.
Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very steep growth before. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues.
Show more