Register and share your invite link to earn from video plays and referrals.

Yo Shavit
@yonashav
ai resilience @foundationOAI. Past: @openai / @HarvardSEAS / @SchmidtFutures / @MIT_CSAIL. Tweets my own; on my head be it.
1.1K Following    10.7K Followers
jfyi Accenture is now TESCREAL
Another thing to note as we’re laughing about this is that Accenture also acquired a Jaan Tallinn funded x-risk “auditing” firm, i.e. the same TESCREAL billionaire Tallinn who led Anthropic’s series A, funds METR & cofounded the Future of Life Institute advising Bernie Sanders 🙃
Show more
kinda fascinated to know whether LTF goes on airwaves to back El-Sayed, do they have it in them
Fascinating to see Republicans run on AI safety. New spot from the group Alliance for A Better Future backing Mike Rogers and using recent comments from Sam Altman and Dario Amodei. Rogers is probably the most vocal guardrails-minded GOP Senate nominee, and now gets air cover for it.
Show more
Happy to sign this. Truly independent evaluation and oversight of frontier AI is a crucial issue. I think one of the biggest bottlenecks is finding people with the technical capabilities and that high bar of independence. AI Eval Forum is a good group to build initial momentum!
Show more
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: Learn more at
Show more
Joe Rogan, welcome to the TESCREAL bundle
Joe Rogan's "crazy" idea for ending all wars: Let AI take over
everyone should chill, it's good to value the natural world and not to hesitate to assert its primacy over the artificial constructs of human civilization
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Show more
0
460
3.6K
431
Forward to community
I wrote this AEF-1 disclosure in METR’s pilot Frontier Risk Report. Transparency about limitations and potential biases felt important to me. My colleagues agreed. If you want the public to trust you to assess AI risks, you try to give them the truth without spin or BS.
Show more
I think it’s underrated that recent cryptography breakthroughs mean we might be able to do internationally-verifiable fully-privacy-preserving inference attestation *with no new hardware or manufacturing*.
Show more
cryptographic evidence of what frontier AI compute is actually doing. If the world chooses to pace AI, @attestable and @prlnet are building the infrastructure to prove it.
Dan Selsam has long been considered one of OpenAI’s most cracked researchers, and I’ve never heard him talk this way before. (He seemed fairly unconcerned before I left.)
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc:
Show more
the apparatchik’s year in year out grind of dull ai governance administrivia will be honored in the world to come
I hope more people will become familiar with the great work of the AI Evaluators Forum, a group including not just METR but also Transluce, RAND, AVERI, Princeton, and SecureBio + others In 2025 they published a standard on independence and transparency covering embedded audits
Show more
FWIW I think @tszzl is wrong here, and it’s strategically important to explicitly divorce the idea of pacing the frontier from restricting open source. The top priority for pacing should be alignment, and open models will be able to leverage the same alignment techniques as closed ones once they catch up to any given capability level. So long as the frontier closed models are paced at the speed of alignment/control progress, and get used widely for hardening/remediation before OS agents at each capability level proliferate across the web, ensuring a continued lagging edge OS ecosystem is plausibly net-positive even beyond Astra level. Politically, I think it’s crucial we narrowly focus pacing on “alignment at the frontier” rather than “misuse restriction”, because frontier alignment covers the most acute LoC risks while costing relatively few parties, and the alternative (misuse restriction) involves costs for many more parties to the point it risks pacing becoming politically intractable.
Show more
@sean_from_earth i won't lie to you, i think open source will be banned before too long after some major disaster. and when the day comes, you'll agree with me. i hope kimi and deepseek etc keep making models but keep them monitored on an api where they should be
Show more
When asked by Henry Kissinger, “what has been the impact of Open Philanthropy’s donation to OpenAI?”, Zhou Enlai famously replied: “It is too early to say.”
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
Show more
I guess we did get a warning shot
ugh unfortunately a banger
Claude-Pop - I'm Upping My P(Doom)
I've known Jacob Coxon for the last three years, having started at OpenAI at the same time. I consider him a deeply thoughtful and measured individual, and also absolute chiller of a guy. I respect his decision to resign from Anthropic, at great financial cost, for his beliefs.
Show more
0
123
1.7K
103
Forward to community
"THERE'S NO WAY I ALONE CAN MAKE A DIFFERENCE. THAT WOULD REQUIRE COLLECTIVE ACTION"
FWIW, the OpenAI and METR reports are great examples of overcoming all of these org frictions to release substantial, potentially-pivotal public scientific evidence. I expect that required major internal will. I've heard people worked >14-hour days to make it happen. They should feel intensely proud. I hope they see how impactful this may be in shifting public preparedness, and want to keep leaning into it. There will be more areas that will require expansive public evidence-sharing if we are to make it through.
Show more
*We need all the scientific evidence on severe misalignment out in the open, urgently.* TL;DR If severe misalignment is indeed emerging absent a high bar of alignment+control execution, the top priority actions for OpenAI and Anthropic in the next couple months should be to make public all necessary evidence for non-safety AI technical experts to become convinced of the science around loss-of-control themselves. [the below is all my personal opinion and doesn't represent my employer] Any meaningful solution to prevent loss-of-control-of-AI will require pan-US-AI-industry adherence to costly practices around AI control, alignment, and security. Crucially, this will include Meta, SpaceXAI, and even NVIDIA, and almost certainly require the blessing or coordination of the US government if it is to be internationalized. The current politics of AI make it unlikely that even OpenAI and Anthropic's joint technical consensus would be sufficient to motivate USG action so long as the lagging AI players are against it. Even if OpenAI and Anthropic were to go to the White House today with slam-dunk technical evidence of an imminent national security threat of AI loss-of-control absent intervention, the USG would probably not trust in their own in-house reasoning about that evidence, and would look to the views of the other admin-friendly AI players and advisors. These other actors appear largely skeptical. The only solution is science.
Show more
@OREAXEAX I think Anthropic keeps doing CoT pressure but I don’t know that OAI does?