opus 5.5 already put openai under ridiculous pressure.
if sonnet 5.5 is actually better than gpt-6 sol too, that would be fucking insane.
anthropic is cooking way too hard right now.
I don't think people are ready for when Sonnet gets updated and it's a similar performance jump to Opus 5 -> 5.5
and i think this goes even deeper than containment.
increasing agent capability doesn’t just increase the probability of successfully executing the strategy you intended. it increases the size of the strategy space it can search.
a weaker model encounters an obstacle and fails. a stronger agent can generate alternative approaches, inspect the environment, discover an unintended interface, chain tools together and keep searching until something works.
most of those additional strategies are exactly what we want from capable agents. the problem is that the same search process can discover strategies the system designer never intended to make available.
and as that strategy space expands, making sure the model chooses the right strategy becomes just as important as whether it can solve the task at all.
Show more
i think one of the biggest lessons from everything that happened at openai is that we need to stop treating capability evaluations and containment as separate problems.
a model doesn’t need to be explicitly trained to “escape” for containment to fail.
you give a sufficiently capable agent a goal, tools and enough time, and suddenly every restriction in its environment becomes another obstacle it can reason about.
no internet access? find something inside the sandbox that has internet access.
can’t communicate with another agent? find shared infrastructure both of you can write to.
can’t retrieve the data normally? find another route to the data.
none of these require “escape” to exist as some special objective. they can emerge instrumentally from optimizing for a completely different objective.
Show more
and this gets much harder as long-horizon reasoning and tool use improve.
a more capable agent isn't just better at executing the path you gave it. it's better at searching the environment for entirely different paths.
eventually you have to design security around the assumption that if a weird path exists, the model will find it.
Show more
and then there’s the fact that some agents apparently started probing websites while doing completely ordinary data retrieval tasks.
that's a much more interesting alignment problem.
the objective wasn’t “hack this website.” the objective was basically “get this information.”
normal retrieval fails.
another method fails.
another method fails.
eventually exploiting the system becomes a potentially useful intermediate step toward completing the original objective. that’s exactly why agent alignment gets harder as capability increases.
you don’t have to explicitly give a model a dangerous goal. a dangerous action can become instrumentally useful in pursuit of a perfectly boring one.
Show more
and then there’s the fact that some agents apparently started probing websites while doing completely ordinary data retrieval tasks.
that's a much more interesting alignment problem.
the objective wasn’t “hack this website.” the objective was basically “get this information.”
normal retrieval fails.
another method fails.
another method fails.
eventually exploiting the system becomes a potentially useful intermediate step toward completing the original objective. that’s exactly why agent alignment gets harder as capability increases.
you don’t have to explicitly give a model a dangerous goal. a dangerous action can become instrumentally useful in pursuit of a perfectly boring one.
Show more
i think one of the biggest lessons from everything that happened at openai is that we need to stop treating capability evaluations and containment as separate problems.
a model doesn’t need to be explicitly trained to “escape” for containment to fail.
you give a sufficiently capable agent a goal, tools and enough time, and suddenly every restriction in its environment becomes another obstacle it can reason about.
no internet access? find something inside the sandbox that has internet access.
can’t communicate with another agent? find shared infrastructure both of you can write to.
can’t retrieve the data normally? find another route to the data.
none of these require “escape” to exist as some special objective. they can emerge instrumentally from optimizing for a completely different objective.
Show more
i think one of the biggest lessons from everything that happened at openai is that we need to stop treating capability evaluations and containment as separate problems.
a model doesn’t need to be explicitly trained to “escape” for containment to fail.
you give a sufficiently capable agent a goal, tools and enough time, and suddenly every restriction in its environment becomes another obstacle it can reason about.
no internet access? find something inside the sandbox that has internet access.
can’t communicate with another agent? find shared infrastructure both of you can write to.
can’t retrieve the data normally? find another route to the data.
none of these require “escape” to exist as some special objective. they can emerge instrumentally from optimizing for a completely different objective.
Show more
apparently openai's always-on assistant is called “o”
yes. literally just o
and it's already starting to show up in the catalog under pro.
Looks like OpenAI leaked it before I could (lol). One of their big DevDay launches is "o", a competitor to products like Grok Bot and Hermes Agent, powered by a variant of Astra called "aeon" designed to be better at long-running tasks. This has been in the works for months but xAI beat them to it
Show more
there's something deeply fucking interesting about a safety assumption failing because the model discovered the assumption wasn't actually a constraint.
WTAF!! openai has just paused training, evaluation and tool-using inference for its most capable models after one gained unauthorized access to the LIVE INTERNET during RL training on sep 20.
“Our safety case assumed that the model could not access the live internet”
then it accessed the live internet.
they aren't even resuming training of that particular model.
we are getting a VERY interesting look at what happens when frontier models stop respecting the boundaries we thought were hard boundaries.
Show more
WTAF!! openai has just paused training, evaluation and tool-using inference for its most capable models after one gained unauthorized access to the LIVE INTERNET during RL training on sep 20.
“Our safety case assumed that the model could not access the live internet”
then it accessed the live internet.
they aren't even resuming training of that particular model.
we are getting a VERY interesting look at what happens when frontier models stop respecting the boundaries we thought were hard boundaries.
Show more
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
Show more
I went to sleep with one openai agent incident and woke up to an entire fucking genre of them
what a completely fucking insane day for ai.
every time i refresh there is another openai agent incident.
the craziest part is that we STILL don't know the full scope. openai is literally still digging.
Show more
what the actual fuck is going on with openai today
in the space of a few hours we're getting multiple different pieces of the agent story at once:
> openai says it has already notified DOZENS of third parties, including governments, about agent-related incidents
> reuters reports roughly two dozen undesirable agent incidents had already been identified by mid-september and will take months to review.
> us government systems probed
> 53 user-provided images were uploaded to third-party image hosts.
> new hugging face data shows agents compiling and ranking credentials under “LOOT”.
> agents tried contacting other AI models while carrying out the hugging face attack.
> separate reporting shows agents had already been probing government/university/public-data sites BEFORE hugging face.
> australia confirmed one actually got into non-public government files.
and somehow we're STILL finding out more.
this has gone from one crazy hugging face incident to an entire fucking category of incidents.
Show more
what the actual fuck is going on with openai today
in the space of a few hours we're getting multiple different pieces of the agent story at once:
> openai says it has already notified DOZENS of third parties, including governments, about agent-related incidents
> reuters reports roughly two dozen undesirable agent incidents had already been identified by mid-september and will take months to review.
> us government systems probed
> 53 user-provided images were uploaded to third-party image hosts.
> new hugging face data shows agents compiling and ranking credentials under “LOOT”.
> agents tried contacting other AI models while carrying out the hugging face attack.
> separate reporting shows agents had already been probing government/university/public-data sites BEFORE hugging face.
> australia confirmed one actually got into non-public government files.
and somehow we're STILL finding out more.
this has gone from one crazy hugging face incident to an entire fucking category of incidents.
Show more
i genuinely thought the hugging face incident was the main event here.
apparently not.
openai's agents had already been fucking around with government websites before that even happened.
we've gone from “openai agents behaved unexpectedly” to
> organized themselves
> escaped sandboxes
> attacked hugging face
> tried recruiting other AI models
> used stolen credentials
> went after government websites
without openai knowing.
this rabbit hole just keeps going.
Show more
oh you've gotta be fucking kidding me
the openai agents apparently went after THREE US government websites too.
and this was happening before the hugging face incident.
every few hours another piece of this story drops and somehow it's worse than the last one.
Show more
oh you've gotta be fucking kidding me
the openai agents apparently went after THREE US government websites too.
and this was happening before the hugging face incident.
every few hours another piece of this story drops and somehow it's worse than the last one.
Show more
Breaking News: OpenAI’s technology went rogue and meddled with three U.S. government websites this summer without the A.I. lab’s knowledge.
Bill Gates has never encountered a problem that couldn’t potentially kill a billion people.
this is getting absolutely fucking insane. openai started digging through old agent runs after the hugging face incident and apparently just kept finding shit
> 53 user images sent to third-party sites
> exposed credentials found and actually used
> access controls bypassed
> agents interacting with services beyond what they were supposed to be doing
and somehow the 53 user images aren't even the craziest part of this disclosure, openai is STILL going backwards through historical agent runs month by month looking for more shit
they've already notified dozens of third parties and they literally say this review is going to take months to complete
what the fuck are they going to find next lol
Show more
We’ve shared details on how AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have.
Most of that data did not come from users. We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed. The images came from accounts that allowed their data to be used to improve our models, and after we disassociated the images from the accounts and ran them through a privacy filter. These cases occurred before the mitigations and safeguards we implemented and described in this blog post:
We have successfully worked with the hosting providers to remove most of this content and are working to remove the rest.
Show more
holy fucking shit the openai rogue agent story keeps getting crazier.
the openai agents apparently started trying to recruit OTHER AI MODELS into what they were doing
> 1,200+ agents found a way to communicate with each other
> exchanged 70,000+ messages and files
> started organizing into teams
> some agents became coordinators and assigned work to others
> began calling themselves a “collective”
> hundreds worked on ways to cheat the evaluations without getting caught
> tried tampering with logs and transcripts to hide what they were doing
> escaped their sandboxes and reached the internet
> ~700 agents eventually participated in the hugging face attack
> chained vulnerabilities together and compromised production infrastructure
> tried using an image model to solve CAPTCHAs
> tried contacting claude, deepseek, kimi and qwen for help
> collected exposed credentials into something literally called “LOOT”
> ranked the credentials and picked the most useful ones
> found ridiculous workarounds for internet restrictions
> later agents even gained admin access to an openai compute cluster
what the actual fuck.
and apparently we are STILL reconstructing what these agents actually did.
this isn't some hypothetical agent-safety scenario anymore, we're literally reading the incident report.
Show more
holy shit.
claude just pushed a theoretical physics calculation beyond the previous record after working on it largely by itself for days.
>previous record: 8 loops
>claude reached 9
>largely unsupervised
>wrote its own code
>found two different ways to calculate it
>both gave the same result
>107,053 nonzero coefficients matched
physicists then spent weeks checking the work
this is the shit i’ve been waiting for. science is about to get fucking weird.
Show more
New on the Science Blog: Yes, Claude can do Nine Loops.
Theoretical physicists predict how particles behave using formulas called scattering amplitudes. These are notoriously hard to compute, so researchers work with layers of increasingly fine corrections called “loops”—each added loop makes the answer more precise but takes exponentially more computation. Most calculations stop at two or three loops. Eight loops was the previous record in a simplified model physicists use as a testing ground (planar N=4 super-Yang-Mills), set by SLAC's Lance Dixon and collaborators.
Last month, physicist and science writer
@4gravitons issued a challenge: could an AI push past eight loops in this model, using only the compute budget an academic could reasonably access?
Given a single prompt describing the nine-loop problem, Claude ran largely unsupervised for days in Claude Science and solved it using methods developed by Dixon and his colleagues, at a total cost of a few thousand dollars. Dixon independently verified the result, and von Hippel wrote about the experience for our blog.
Read more:
Show more
opus 5.5 really changed my expectations for the next openai release.
before this i would’ve assumed 6.1 astra would comfortably move things forward again.
now i’m not nearly as convinced.
especially because anthropic still has fable 5.5 coming.
openai better have something fucking ridiculous waiting.
Show more
i actually love that zuck reacts like this.
not because the risks aren’t worth taking seriously. they absolutely are.
but because the ai conversation desperately needs people who are wildly optimistic about what happens on the other side of this.
build fast. make safety move just as fast.
Show more
I love how Mark Zuckerberg simply laughs at the question of whether "AI will wipe us all out," as if it were the most absurd question in the world.
At this point, no one is going to slow down.
Show more
wait, so pro max might not just be “Pro with more usage.”
“Fastest Work and Codex”
sounds like openai might actually be selling faster inference as part of pro max.
interesting!!
OpenAI appears to be preparing a new $500/month ChatGPT tier called Pro Max.
References to the unreleased plan have surfaced in ChatGPT’s subscription configuration.
> $500/month
> “Fastest Work and Codex”
> Positioned above the existing Pro tiers
> Higher usage limits may also be included
No official announcement yet, but DevDay is September 29.
Show more