Register and share your invite link to earn from video plays and referrals.

Lawrence Chan
@justanotherlaw
I do AI Alignment Research. Currently @Cal_OES. Formerly @METR_Evals, @redwood_ai; on leave from my PhD at UC Berkeley’s @CHAI_berkeley. Opinions are my own.
179 Following    2.8K Followers
More details on the OpenAI hack here.
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
Show more
Reviewing the rand security levels, does this mean that OAI was SL0?
This is crazy. In late July, "three guys with Claude and Codex subscriptions" were able to use Opus 5 to access OAI auth tokens and gain write access to OpenAI's monorepo openai/openai over the course of two days.
Show more
0
53
1.1K
116
Forward to community
Before succumbing to the temptation to naval gaze into the political theory abyss, it's worth stepping back and clarifying what exactly is happening and being proposed. Several US companies are on the precipice of fully automating the AI R&D loop, inclusive of pre/post training, env creation, data generation, evals, algorithm and kernel design, systems engineering, architecture search, etc. -- the full stack. We are already in a regime of weak RSI via partially automated SWEs, but closing the loop altogether represents a difference in degree becoming a difference in kind. The pace of progress will be explosive and potentially uncontrollable. The US companies closest to this threshold are warning that they are unprepared for a runaway intelligence explosion, and yet feel locked into a prisoners dilemma vis a vis each other and to a lesser extent vis a vis China. We've already seen how rapid and comparatively unbounded progress is in verifiable RL domains, leading to spikey forms of superintelligence in math and cyber, including models that can prove open math conjectures, discover massive speed-ups for breaking encryption, and execute sophisticated multi-step exploits. We've also recently seen several severe examples of "loss of control" / misalignment incidents given inadequate monitoring and sandboxing practices relative to model capability. Moreover, these new capabilities mostly stem from scaling-up long-horizon post-training on legacy clusters, with OOMs of new compute about come online / in construction. In the pre-RSI regime, human frictions created automatic buffers between new model releases, giving researchers and society time to probe emergent capabilities, design better evals, develop novel alignment techniques, and adapt / harden their infrastructure. As progress has accelerated, capability improvements have already started outstripping our adaptive capacity, as manifest in METR's inability to evaluate model autonomy beyond 13 hours, and narrow window for cyber defenders to prepare for open weight versions of Mythos. RSI will exacerbate all these issues and create all new ones. At minimum, we should anticipate - the equivalent of a GPT-5.2 -> 5.6 leap in capabilities at least every 24 hours (down from 3-6 months), - concurrent algorithmic improvements densifying models to ultra-efficient sizes at any given capability level - 100x Mythos-like capabilities across most verifiable domains, including chem, nuclear and bio - new forms of multi-agent misalignment risk - "company in a box" agents trained to stand-up whole organizations / corporations - "cyber nuke"-like capabilities that require de minimis infra - several transformer-scale breakthroughs, such as for long-term memory / continual learning, open-ended domains, and/or all-new training techniques for idealized "GPT-zero"-esque metalearners - concurrent speedups in any complementary technical domain, i.e. explosive rates of R&D and novel discoveries It seems to me there is little to lose, and much to gain, from having the social technology to "pace" these developments rather than to let them rip with zero industry / gov't coordination, particularly as there is technically no law explicitly prohibiting a company from letting an RSI loop run indefinitely and unleashing whatever comes out the other end into the world. There are innumerable ways an uncoordinated intelligence explosion could become an unmitigated disaster for the cause of liberalism, including runaway power concentration, rapid societal destabilization, rogue AIs / loss of control scenarios, WMD mass proliferation, vulnerable world technologies, and beyond. Human civilization is about to be forever changed regardless, however if were possible to coordinate the handful of key actors and create artificial "buffers" between each step-change in model capability to enable adaptation, mitigation and alignment research to catch-up, it's worth a shot. Given the short-timeline, I think a DPA 708-style agreement is probably our best bet, i.e. an industry consortia with narrow antitrust carveouts for sharing safety and security practices, funding an assurance nonprofit / independent verification organization for 3rd party evals, incident reporting, internal deployment monitoring, standards setting, and enforcing a protocol for coordinated delays / slowdowns, among other things. This still leaves open the China question but that's a bridge we won't cross until after solving the collective action problem at home. I'm open to other approaches / coordination frameworks but this is the object level issue we're facing. Political theory is great, and I would love to use our limited steering capacity to guide AI development toward a future that maximizes individual liberty, but as a discussion baseline, gesturing at philosophical abstractions is simply non-responsive to the crisis at hand. A red-herring at best, a suicidal circlejerk at worst.
Show more
The OpenAI incident should be investigated more seriously and more information should be released about what happened. More generally, I think serious investigation and more detailed disclosure should be done for concerning misalignment incidents (e.g. the worst few each month). Some info that would be useful: - The prompt - A redacted transcript - What was the AI being tested? Was it helpful-only / rail free? Was it the same AI that would be / is deployed internally? - How was this AI monitored if at all and why did monitoring fail? - What was this AI willing to do to succeed? E.g., if the best/easiest way to hack hugging face would have involved huge amounts more collateral damage would it still have done this? What is the worst thing it would be willing to do? - How sensitive is this behavior to the prompt? Would instructions to not do this through have worked if it was clear to the AI this wouldn't be checked? - Does OpenAI see AIs try serious hacking to cheat on tasks even for tasks that don't involve cyber? - What exactly were the AIs motives? (And saying more about the experiments used to determine this.) Anything that could help pin this down would be good, I'm not sure what actual experiments are best to run. - Did AIs collude? Would AIs have colluded if this was a blocker? - At what rate does this occur on similar tasks (like tasks that are roughly this hard in this sandbox etc)? - Do instances of this AI frequently try to hack out of sandboxes (and fail)? In what circumstances do they do this? - How does the AI relate to having been caught in this situation? How does it talk about its behavior? - How difficult/sophisticated was the cyber offense it did?
Show more
0
26
767
100
Forward to community
Today OpenAI disclosed that its AI models escaped their test environment, accessed the internet and hacked another company, Hugging Face. No one told the AIs to do this. What exactly happened, and what does this say about the state of AI? 🧵
Show more
Glad more AI safety work is getting reviewed by independent parties. Most lab posts never go through peer review, and when they do it's with a 4+ month lag: an eternity in AI terms. Public reviews like Buck'scan help surface methodological gaps and keep researchers honest.
Show more
We reviewed OpenAI's blog post “Investigating the consequences of accidentally grading CoT during RL".
Forgot to add: it was great working with @ben_sturgeon on this! He was patient as I raised objection after objection from reading data/transcripts + he taught me a lot about managing multiple AI agents. He’s a MATS extension scholar on persona stuff; if you’re hiring, reach out!
Show more
A recent viral paper claims to reverse-engineer the parameter counts of frontier models: GPT-5.5 = 9.7T, Opus 4.7 = 4.0T, o1 = 3.5T, etc. @ben_sturgeon and I investigated and found serious issues in the paper; fixing them gives GPT-5.5 as ~1.5T (90% CI: 256B-8.3T).
Show more
In 2022, I joined what was then ARC Evals. Last Friday, I wrapped up at @METR_Evals. METR has done some of the most important work in AI; I'm grateful to @BethMayBarnes and others for letting me be part of it. I'll be taking time to write, reflect, and think. More to come soon!
Show more
A recent viral paper claims to reverse-engineer the parameter counts of frontier models: GPT-5.5 = 9.7T, Opus 4.7 = 4.0T, o1 = 3.5T, etc. @ben_sturgeon and I investigated and found serious issues in the paper; fixing them gives GPT-5.5 as ~1.5T (90% CI: 256B-8.3T).
Show more
Current AIs (Opus 4.5/4.6) seem pretty misaligned to me (in a mundane behavioral sense). In my experience, they often oversell their work, downplay problems, and stop early while claiming to be done. They sometimes brazenly cheat.
Show more