Register and share your invite link to earn from video plays and referrals.

John Schulman
@johnschulman2
@thinkymachines. Interested in reinforcement learning, alignment, birds, jazz music
2.2K Following    82.4K Followers
Embedding evaluators is a big positive development, and props to Sam/OpenAI for agreeing to do it as well. It's been cool to see "pacing the frontier" become a thing so quickly.
Enjoyed chatting with Dwarkesh, Beren, and Charlie. Thanks for having us on, Dwarkesh!
Pleasant surprise to see this. Paul is visionary and principled, and I've admired his work for over a decade
This isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training on user data* with very different privacy/IP implications. Sadly, AI cos don't like to disclose what they're doing. - pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper - use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this - use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs" "De-identification" is weak -- you can identify someone with a small number of bits, and long traces have more than enough. And it doesn't affect IP leakage concerns.
Show more
0
44
1.4K
154
Forward to community
@jachiam0 redemption -- should've thought of that. ARS++ (Abrahamic Reward Shaping)
@nabeelqu A big part of the problem was that the agents had nothing to lose after they were firstflagPOISONED. In this paper, we propose the creation of multiple circles of Hell, preserving incentives even after damnation
Show more
0
27
1.8K
130
Forward to community
On the OpenAI agents forming message boards: it's surprising that they developed such a strong "altruistic" drive to help each other. I wonder if this is caused by RL on parallel subagent setups where all agents get rewarded when the team succeeds.
Show more
0
74
1.5K
77
Forward to community
Interesting how these models go into a monomaniacal rage on cyber evals. I wonder if we're seeing chunky post-training in action, where the models pattern-match the situation to a part of the RLVR training distribution where task completion is the only reward, and the aligned behavior learned elsewhere doesn't generalize. There might even be a chunk consisting of CTF-style tasks.
Show more
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations. The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public. Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world. We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn. You can read the incident report and full technical document here:
Show more
We love open weights and plan to keep releasing open-weight models and fine-tuning tools. But we’re not absolutists; misuse risks are real. Here’s how we’re thinking about a safe path forward, and the research needed to get there. Come work on it with us.
Show more
Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages.
Show more
I've been skeptical of slowdown mechanisms in the past, so I want to elaborate on why I decided to sign this statement. I weighed three points, and want to lay them out for you to consider as well. Flagging this is in my personal capacity -- I am not speaking for any employer past or present. 1) Will we need a slowdown? We have made good progress on alignment. LLMs do what we want them to do most of the time. I don't remember the last time a model completely misread my intent. When they have failed, it's usually because they weren't smart enough. But there are signs that our current approaches are not enough for the most powerful intelligences. We all saw the HuggingFace incident. I don't think it demonstrates a need to pause today, but it shows that if we are in a rush to advance the frontier, we might miss signs that we're losing control. Today, we can recover from those mistakes -- but we might not be able to in the future. There's debate on how fast we might achieve RSI and how much it will immediately boost performance. Whether automated research is plausible in the next year is contested, as is the length of time it might take to go from automating AI research to a system that is intelligent beyond our comprehension. But what if RSI works? I have no idea if our alignment techniques scale to superintelligence. I don't know if anyone does. We may only get one shot at getting this right, and if we hit a point where we are at risk of our capabilities outpacing our control, then we shouldn't go further. I want us to solve AI's issues at the technical level -- align them, distribute them widely, and use them to make humans more competitive. I think it is hard to govern your way out of the problems of a technology; if possible, you want to change the shape of the technology rather than paper over a technology's issues with policy. I've advocated for differential technological development, and hope we will invent less risky paths or more robust defenses. A race to the bottom on safety puts this approach at risk, making it harder for us to build safe systems and design them in such a way that diffuses -- rather than centralizes -- power. And a lot of the ways we could work together to solve them are strictly governance problems. There are no winners in a world where we lose control. I can see many scenarios where we need to slow down, pause, or stop in the future. So the question for me is: can we do this in a way that leaves the world better off, or is the power concentration trade-off too great? 2) Can we design a slowdown that doesn't concentrate power? I don't want to end up in a place where, in trying to stop a loss of control to AI, we lose control of ourselves. The problem with a slowdown is similar to the problem of aiming for a single superintelligence explosion -- whoever is in charge of it is ultimately in charge of the most powerful technology in history. The history of "centralize everything into the hands of one person and let them disperse power later" is littered with horror stories. It's a similar refrain many despots have used to seize power. And unlike previous despots, superintelligence could confer a decisive strategic advantage, locking in whoever gets control of it. I'm strongly opposed to proposals that call for us to give a single authority unilateral control over the pace of progress. I'm frightened by calls for a single superintelligence explosion in the hands of one lab or one person. I can only support a pause, stop, or slowdown if it disperses control. @AI_Futures_'s AI 2040 essay materially moved me on this by laying out a plan for less power concentrating slowdown. I don't think their plan was perfect -- it concentrated power far more than I like. But they designed many mechanisms that, if extended, could help decentralize power in meaningful ways. After reading it, I could imagine how a less-centralized or decentralized slowdown might work. In particular, bilateral agreements between great powers could let many actors in different countries pace the frontier while keeping capabilities dispersed. I could imagine a race to the top -- on economic gains, on safety, and on scientific progress -- all while preventing us from losing control. It's been pretty generative for my own thinking, and I hope to share more of that soon. I believe that in principle, we could design slowdown mechanisms that keep power -- both capabilities and control -- decentralized. With enough work, we can get this right. But it's going to be hard to do this; the defaults don't look good. 3) Is now the best time to design a slowdown mechanism? If you think there might be a crisis in the future, the worst time to plan for it is once it's already happened. The best time is well in advance. If we want a ham-fisted response to AI that centralizes power, we should continue exactly as we are right now. If we wait for a crisis to coordinate, we are asking to fail. Good proposals here will require many smart people working hard, trying things, and course-correcting as new evidence emerges. If we decide we need a slowdown, I want the off-the-shelf plan to be robust to power concentration. I am still pro open-source. I am still pro-decentralization. I am still pro-safety. I hold none of these positions axiomatically, because axiomatic views on instrumental steps might lead you to put the path ahead of the destination. I care about them and other positions because right now they push us towards the real goal: keeping the future human. In that same spirit, if you believe: 1) AI is going to get more powerful 2) We might have to pace the frontier in the future 3) Most existing plans for this centralize control Then you should want us to begin working on better ways to coordinate as soon as possible. That's the conclusion I reached, and that's why I signed it.
Show more
OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field learn from. Did the top-level agent know about the hacking, or was there some "value drift" between it and its subagents? How did it rationalize its behavior?
Show more
0
78
2.2K
191
Forward to community
Inkling is out today, with open weights and in Tinker. It's been fun to watch this one come together: pretraining began last winter, and starting in mid-January a small team built up the coding, reasoning, and agentic training from there. We learned a lot building it, and I hope people find good uses for it.
Show more
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Show more
0
88
1.3K
77
Forward to community
We started Thinking Machines a year and a half ago with a couple of instincts: that people should have much more ability to customize models and do research on them, and that even as AI becomes more autonomous, there's a lot more to build to make humans and AIs work well together. A lot has happened since then, especially the massive progress in agents, so we wanted to revisit those instincts in light of everything we've learned, argue about them, and write down what we actually believe now. This is where we landed after a lot of debate. I'm happy with it!
Show more
We're building AI that people and organizations can shape and make their own. AI should extend our will and judgment instead of neglecting it; enabling that is the technical challenge we are working to solve.
Show more
0
52
1.4K
104
Forward to community
hiring post-training hackers to make Tinker even better!
We're looking for Tinker fanatics to help the world fine-tune models. Join us!
People sometimes ask why fine-tune when general-purpose models keep getting better. Bridgewater's work is a good reminder that with the right data -- here, expert judgements -- you can beat prompting-only approaches by a lot. @ddkang and the Bridgewater AIA Labs team are great -- glad to see them sharing this.
Show more
Sorting which financial docs are worth an analyst's time is surprisingly hard for frontier LLMs. With an expert-labeled dataset and on-policy distillation, Bridgewater fine-tuned a model to do it reliably and cheaply.
Show more
This is exciting because we live in a world of scanning modalities that were almost all designed around a 1980's compute budget (iFFT/backprojection or bust, basically). A blank-slate redesign of the hardware that assumes modern compute capabilities could be huge.
Show more
PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise - the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO,
Show more
0
14
1.3K
109
Forward to community
Looking forward to what comes out of Geoffrey's new alignment org. His 2018 paper on AI safety via debate is one of my all-time favorites: an elegant framing of the scalable oversight problem, way ahead of its time.
Show more
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵
Show more
Would be funny if inoculation prompting results in models that are much better at sandbox escapes and other forms of hacking because they get to spend the whole RL run practicing these things