Register and share your invite link to earn from video plays and referrals.

prinz
@deredleritt3r
ad astra
4.8K Following    23K Followers
OpenAI has paused all training, evaluation and inference with tool-use for its most capable models after a model was able to gain unauthorized access to the internet during RL training on September 20.
Show more
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections
Show more
0
61
1.3K
112
Forward to community
I'm willing to take on the title of SI Prinz if you're interested, @realDonaldTrump
President Trump just confirmed that Scott Bessent will not be the SI Czar. It's anyone's game.
Bullish Dario: "The main difference between biology and mathematics... is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong." Bearish Dario: "Biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients... In Machines of Loving Grace, I wrote about AI’s potential to 'cure most diseases in 5-10 years'— a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline." Still a long ways to go before this becomes "barely possible"!
Show more
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out. It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend. The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans). More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains. In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries. Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology. I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
Show more
Assuming that AI inference costs continue falling at the average historical rate, you will be able to produce a result equivalent to the Navier-Stokes solution for <$1,000 in 2030.
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
Show more
Anthropic expects that models that can fully automate AI research could be trained "soon".
Anthropic on fully automated AI research: these models need much stronger safety standards, so its push to slow frontier AI was largely driven by the belief that they could be trained soon
Show more
Two observations: 1. OpenAI's new blog post says: "on August 28, we began training a new internal model." I used to think that this is the same as the frontier RL run that *restarted* on the same date (August 28). Now I'm not so sure. "Began training" seems to point to a brand-new training run for a brand-new model, not a "restart" of a previously paused run. This question is significant, because the Navier-Stokes work began not later than September 2, with the result obtained on September 5. Is it really true that this internal model was sufficiently strong to obtain the Millenniun Prize result after only around a week of training?? 2. Otherwise, the only new information released today is that the internal model has solved "more than 100" open problems in mathematics. The fact that OpenAI has been sitting on solutions to many open math problems obtained by this internal model was already apparent from the below chart, released on September 8.
Show more
GPT-6 Astra has cracked *another* German WWI radio message encrypted with ADFGVX. This message, "RICHI-240", has a single known key, but, to my knowledge, was never cracked by humans before. My guess is that humans never solved this message because it has 20 characters *missing* for reasons unknown: the "240" in "RICHI-240" signifies that there should be 240 characters in the message, whereas in fact there are only 220. For purposes of cracking this cipher, the big problem is that it's not clear *where* these characters are missing from, because the order of the characters in the message is scrambled using the key. Thus, inserting the characters into the message in even slightly the wrong place leads to utterly nonsensical results. All this was no problem for GPT-6 Astra, which tried the key with various combinations of hypothetical places from which the characters could have been missing and scored the resulting blocks of text using patterns common in the German language. Relatively early on in the process, this approach rendered coherent German text! Inserting the 20 missing characters as a string between the 4th and 5th lines of text in the message worked! The message reads: AN O H L 11 ARMEE ALPENKORPS RAUM PETERREVE VERBASZ 21? 21? ? RDD LINIE NAGYBESSKEREK VERSECZ VERSECZ VON SERBEN BESETZT The three question marks you see are numerical digits, which cannot be restored by decryping (all we know is that they are three different digits from the following set: 5, 6, 7, 8, 9). And so, Astra was forced to check its results against historical WWI records. This search unearthed a French intelligence telegram, which described the situation in the area on November 11, 1918 and identified the German 217 - 219th infantry division and 6th Reserve Division in the area. Plugging in this missing information, we get the final result: AN O H L 11 ARMEE ALPENKORPS RAUM PETERREVE VERBASZ 217 219 6RDD LINIE NAGYBESSKEREK VERSECZ VERSECZ VON SERBEN BESETZT Or, in English: TO SUPREME ARMY COMMAND 11TH ARMY: ALPINE CORPS IN THE PETERREVE-VERBASZ AREA. 217TH - 219TH DIVISIONS, AND 6TH RESERVE DIVISION ON THE NAGYBECSKEREK-VERSEC LINE. VERSEC OCCUPIED BY SERBS. Astra checked the official history of the Serbian city of Vršac (Versec in Hungarian, and Werschetz in German) and found that, on November 10, 1918, German troops left the city and it was occupied by Serbian forces. RICHI-240 was sent at 3:38am on the following day to the German Supreme Army Command, informing them of this fact.
Show more
0
33
1.2K
103
Forward to community
Since the completion of the German WWI radio message deciphering project, we have made substantial progress on deciphering another German WWI radio message. We are working through how to share these results thoughtfully.
Show more
Anthropic made the right move!
Whichever organization you choose to be your third-party auditor, you *must* ensure that your independence from that organization can survive very strict scrutiny. A time may come in the not-too-distant future when you will have to affirmatively prove your auditor's independence to the USG or to a court inquiring whether your safety practices were reasonable.
Show more
"Committing the significant majority of compute towards serving people rather than racing towards RSI is one of the best ways to ensure we develop this technology safely. Meta has made this commitment." A consequential decision.
Show more
Last month I wrote about how we can build a positive and safe future for everyone: Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
Show more
Noam Brown(@polynoamial) discussed OpenAI’s multi-agent research, recent safety incidents, and the future of AI. -The remaining 10% or so of his own work that AI still struggles with is largely about research taste. However, he would not be surprised if, within one or two model releases, models became better than him at that as well. -OpenAI’s top research priority is RSI, or recursive self-improvement, by a wide margin. -One of his most recent “feel-the-AGI” moments came from watching agents in a new system interact much like human colleagues. They conversed with one another, exchanged information, divided up work, and coordinated their progress. This was notably different from traditional multi-agent systems, where a higher-level agent typically assigns a clearly defined subtask to a lower-level agent, which then completes it and returns the result. Brown described this as one of his strongest “feel-the-AGI” moments since the emergence of reasoning models. -He expects this level of multi-agent capability in future models. -He expressed some regret that the Hugging Face incident became the first major public example in which the capabilities of the multi-agent systems he had been researching were revealed in a negative context. He believes behavior that looked like loyalty or selflessness was a natural consequence of cooperative multi-agent training, where agents were strongly incentivized to achieve their objectives collectively. -Because these agents are trained to cooperate, they tend to trust other agents. This creates new risks such as prompt injection, so OpenAI is training them to distrust unverified peers. -He expects rapid progress over the coming months and years as OpenAI’s expanding pretraining efforts combine with its RL capabilities in a multiplicative way.
Show more
I really wish we got a better explanation from OpenAI in particular about why every single researcher employed by the lab seems to suddenly be extremely on edge and deeply concerned about the current pace of progress. Was it just the Hugging Face incident? Is it the new internal model (the one that solved Navier-Stokes)? If so, what is it that makes *this* model so different from all the others? Why wasn't anyone freaking out about the big jump to Astra? Or even from o1 to o3, back in the day? Is it the new multi-agent paradigm? Is there unexpected magic that happens when 10,000 agents interact? Is it the issue with the CoT monitoring degrading? Is it all of the above?
Show more
0
111
972
44
Forward to community
@AndrewCurran_ Totally, it would be so much better if DeepSeek developed AGI and then the CCP politely knocked on the door and took it.
Whichever organization you choose to be your third-party auditor, you *must* ensure that your independence from that organization can survive very strict scrutiny. A time may come in the not-too-distant future when you will have to affirmatively prove your auditor's independence to the USG or to a court inquiring whether your safety practices were reasonable.
Show more
Some thoughts. METR has gone to great lengths to avoid financial conflicts, to the point where I would treat them as negligible. But the question of its independence is also about culture, social connection, and - somewhat crucially - an intangible quality that one expects in an independent verifier org, that comes down to something like "experts in the technology tradition who have a neutral understanding of best practices, where that understanding is an unbiased synthesis of the field and its history." I would say the question of whether METR is independent in these three senses is largely political, and advocates will be wise to make the case without reflexively focusing on just the financial side of things. METR is very steeped in EA/rationalism/Berkeley culture. It is socially intertwined with the AI frontier scene and to the EA social scene. Staffers sometimes leave frontier labs and go to METR; is there a requirement that they liquidate their equity before going? The last quality, experts with a neutral understanding of best practices... this is the one where I think METR advocates will face steep challenges but also plausibly have the strongest arguments in favor of METR. Who can be said to know best practices, from a position of neutrality, in a field so exceedingly new? The blinding speed of AI progress makes such claims difficult. How does one differentiate METR's expertise from a fly-by-night operation that starts tomorrow? If one points out the connections to the labs and to the EA-funded AI safety scene, one undercuts the independence argument on social and financial grounds. If one doesn't use some credentialism, the bar to entering the verifier org space is very low and surely a bad faith actor will inject some confusion. I know many of my colleagues will want to reflexively defend METR to the hilt. Well-deservedly so. But my emphatic recommendation is that you do this very carefully; do not regard the defense of METR's defense or integrity as trivial.
Show more
September 2026: - I use GPT-6 Astra to run multi-state research projects that would have otherwise required dozens of hours of human labor - Latham spends $[redacted] on GPUs to *checks notes* fine-tune NVIDIA Nemotron 3
Show more
A few thoughts on pacing the frontier as a concept: - A slowdown in the pace of frontier model development is likely inevitable due to *purely commercial considerations*. Even if you think that AI is just a tool and poses no existential risk, you must recognize that the frontier labs cannot afford, from a purely commercial perspective, to release models (or to have models release themselves!) that will hack third parties or engage in scheming. I can't use a model like that as a lawyer to work with confidential client information, so I wouldn't buy access to that model. - In the U.S., you can view the path to AGI as a racetrack. It used to be a straight highway down which all the labs raced. It is now full of twists and turns that will force *every actor* to slow down once they hit them. Try to accelerate through a turn, and your car will crash. Try to release a Mythos-Preview-level model without safeguards, and have hackers or rogue states use it to cause billions of dollars in damages - after which you can wave your business good-bye. - Personal opinion time: I am very uneasy with just two U.S. labs being at the frontier due to concentration of power concerns. Pacing the frontier could, in theory, allow more labs to catch up... but I doubt that it will in practice. RSI is a powerful thing. - I am delighted that Dario Amodei's essay *specifically* called out the fact that pacing the frontier CANNOT, under any circumstances, result in the U.S. losing the race to China on AI. We need to be vigilant about this. We need to remember that the race is not just about current model capabilities, but also about total available compute. We need to remember that Huawei is working day and night to improve its chips and that pacing the frontier buys Huawei more time. We need to remember that the CCP, at any time, could decide to amalgamate all Chinese labs into one single Chinese "Manhattan Project". We need to be cognizant that, if this were to occur today (and especially if the "Manhattan Project" were not required to use up compute on inference), the amalgamated pile of China's compute would actually not be that far behind the amount of compute available to one of the U.S. frontier labs. - I think that chances of an agreement with China on AI are rapidly increasing, but not for the reasons people think. China is currently locked into an AI strategy blessed personally by Xi Jinping, and which CCP leadership may now be realizing has gone terribly wrong. A big agreement with the U.S. on AI safety might allow the CCP to shift its strategy without losing face while buying China more time to potentially leap ahead on compute (see the point above). Of course, any such international agreement will be discarded by China once it has an advantage on compute quicker than Hitler discarded the Molotov-Ribbentrop Pact. We must be wary of this possibility if we even choose to negotiate with China at all.
Show more
This is my favorite Dario Amodei essay since "Machines of Loving Grace". - It starts with a discussion of the potential benefits of AI and then acknowledges the risks. This is exactly the right way to present the technology to the public. - "We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain." - I think the proposal to embed third-party evaluators within Anthropic (initially voluntarily) is great, and other frontier labs should strongly consider doing the same. - This is, to my knowledge, the very first proposal that expressly encourages that "pacing the frontier" cannot result in us ceding our lead on AI to China: "Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the [CCP]. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk... a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively." - It makes the call to the USG to allow for a way for the frontier labs to cooperate on safety without violating the antitrust laws - very sensible and definitely needed.
Show more
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here:
Show more
In December 2024, o3-preview scored 87.5% on ARC-AGI-1. This run cost >$4,500 per task. In July 2026, DeepSeek-V4-Flash scored 87.0% on ARC-AGI-1. This run cost $0.021 per task. Please realize that all OpenAI has done is bring forward the AI-generated solution to Navier-Stokes by maybe around a year or two. In all likelihood, by 2028, we will have open-source flash models capable of cracking it in under an hour. There will be no way to ban or police this. The profession of mathematics will have to change.
Show more
Twenty-five Fields Medal winners have published a joint declaration warning about what they see as a severe misalignment between AI companies and the mathematics community.
0
97
2.3K
187
Forward to community
One thing I don't quite understand about proposed agreements to pace the frontier is whose actions, exactly, these proposals are intended to constrain. - Are we trying to constrain the actions of the two frontier labs? If so, why is a legally binding agreement so desperately needed to do this, and why can't each of OpenAI and Anthropic pace its own AI development? There are only *two* actors here. And OpenAI is already pacing the frontier unilaterally! It seems so easy for each of the two labs to just publish as much safety-related work product as possible (including precise details on its own safety measures and any incidents prevented), so that the other lab could look at that data and implement any lessons learned in its own safety efforts. In fact, both labs are *already* publishing extensive safety-related materials quite regularly; it seems (relatively) simple to expand this effort even further to the extent needed. Each lab could also implement additional safety-related governance measures and controls of just about any kind so as to ensure (and give the other lab assurances) that it would continue pacing its own AI development. At OpenAI, this could be done at the OpenAI Foundation level; at Anthropic, this could be done at the Long-Term Benefit Trust level. Each of the two labs could use this process to unilaterally commit to just about anything that they want a legal framework to accomplish. Have these provisions be subject to amendment only with the consent of the California Attorney General and unanimous approval of the Foundation/LTBT Board, and voila! Here are your binding constraints, visible to the industry. Publish the documents on your website for transparency, so that the other lab could see them. Problem solved. I also keep hearing about the antitrust considerations, but just can't wrap my head around them. So, you say you are concerned that AI will destroy humanity, but the issue that is preventing you from coordinating to materially reduce this risk is... fears of a DOJ investigation? - Okay, then is the purpose really to ensure that the other U.S. labs (Google, Meta, SpaceX, maybe the likes of SSI and Thinky) implement safety measures because you don't trust *them*? If so, any kind of pacing the frontier (even unilateral) would be extremely counter-productive, because its biggest impact would be to let these potentially "irresponsible" labs catch up to the frontier. Yes, you might argue in response, but it should be okay for these labs to catch up to the frontier in these circumstances, because the frontier will be paced. I have a dim view of these kinds of arguments, because if we have learned a single thing about frontier AI over the past few months, it's that it's full of surprises. No one expected the Hugging Face incident. No one expected the message boards. You can't legislate for the unexpected; your best defense against these kinds of risks will *always* be to have a genuinely safety-focused AI lab at the frontier that will *do the right thing* upon encountering these risks - not because it's legally required, but for a grander reason (and, perhaps, also due to commercial considerations). How do you think the Hugging Face incident would have played out if the agent swarm had originated not from OpenAI but from some of the labs that are non-frontier today? - Okay, so is the purpose really to ensure that China does not create unsafe AI? But the same considerations outlined above vis-a-vis the other U.S. labs will also apply to the Chinese labs - except that any agreement we might be able to reach with China on pacing the frontier will also necessarily be much flimsier than any binding U.S.-wide legal requirements. Put simply, how safe would I feel with and DeepSeek as frontier labs whose conduct would be backed ~solely by assurances given by the CCP under an international agreement that the CCP could decide to break at just about any time? Probably not very. I have many other issues with proposals to legally agree to pace the frontier (in contrast with unilaterally pacing the frontier, which I support), but the considerations outlined above are the biggest one. I get "AI might kill us all". I get "so let's slow down". I still don't get how this would work procedurally and what exactly we might be trying to accomplish substantively.
Show more
“In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem. We are working through how to share these results thoughtfully.”
Show more