A ranty tweet in defence of using numbers to quantify risk:
What might happen if there were no AI forecasts? Would people say "wow we don't know what is going to happen with AI, I guess we should be ambivalent between speeding up and slowing down"?
No! It would probably (in my view) be full speed ahead.
This has forced making poorly justified but precise forecasts to illustrate levels of concern. There are flaws in this, sure, but note that signifcant groups of people in the field don't talk remotely like this about most other things. 1000s of plastic researchers aren't like "there is a 10% chance plastic kills everyone". Oil CEOs are not like "we might end human life". But 1000s of AI researchers do think that. We have polled them, 10% is the median response.
So I don't love P(doom) and I would prefer we were clear that 3+ year forecasts are very inaccurate. But let us not pretend that if people had said "I am very worried" this would have been understood as those people intended. They think there is a plausible imminent risk to everyone alive. Some think the risk is likely, some almost certain.
I am glad they can convey that, using numbers if they have to.
Show more
Today I made the difficult decision to continue to not be recruited by Anthropic or OpenAI. Not being offered equity at this level is not a decision that I make lightly, but I hope it reinforces the sincerity of my convictions for all who deny the act-omission distinction.
Show more
BTW I tried for a while over the last couple weeks to prove a full Lean kernel correct and failed (in the spirit of also mentioning failed attempts). :)
I should make the prediction:
80% that the type theory difficulties in Lean are resolved within a month.
40% that they are resolved within a week.
I should make the prediction:
80% that the type theory difficulties in Lean are resolved within a month.
40% that they are resolved within a week.
Con-leche is safe against ZFC + inaccessibles only via extra checks which are believed unnecessary, but where existing type theory methods don't yet work. Thank you to
@TaliaRinger for emphasizing this! The fastest versions this stuff will exist only once those are solved.
Show more
"THERE'S NO WAY I ALONE CAN MAKE A DIFFERENCE. THAT WOULD REQUIRE COLLECTIVE ACTION"
Con-leche is safe against ZFC + inaccessibles only via extra checks which are believed unnecessary, but where existing type theory methods don't yet work. Thank you to
@TaliaRinger for emphasizing this! The fastest versions this stuff will exist only once those are solved.
Show more
I've heard multiple times that if [lab] where to unilaterally pause, this would accomplish nothing and/or be net negative.
If I understand correctly, _a lot_ of people are reacting to reports of a single employee quitting and a single other employee estimating >10% chance of AI killing all humans.
Imagine if an entire lab paused!
Show more
I think we have a ~50% chance of all dying as a result of superintelligence, mostly due to actions in the next few to 10 years. I don't expect to have that resolved to below 10% or above 90% before we either make it through, or we don't.
We will have to act despite uncertainty.
Show more
This is false. One of the many consequences of superintelligence actually arriving safely would be not making dumb “people who worry about risks are silly if those risks didn’t occur” mistakes.
If there is a real chance of AI loss of control, and it is somehow averted, odds are high the people who made that happen will not get credit and will be remembered, if at all, as silly worry-warts.
Look at how people talked about nuclear risk in the 90s/00s
Show more
Resolution is hiring philosophers to work on AI alignment! Please apply!
The philosophy program at Resolution is growing! If you're motivated to work on helping solve AI alignment, we'd love for you to consider joining us.
Apply here:
Deadline: September 30, 2026 — we'll review applications as they come in.
You don't need a Philosophy PhD to apply. We're building an interdisciplinary team and welcome candidates from different research/professional backgrounds. More info on the research agenda is in the job posting.
Show more
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Show more
Yes, regulation would be great here, but also...you could just tell someone independent what the architecture is?
Framing this as the fault of the people who don't know the details is very weird.
Thank you for $10M in API credits for helping with semiautomated alignment theory research, OpenAI!
It is worth stating explicitly that this is Resolution making a choice to accept funding from AI labs (and lab-adjacent sources in the future). Which is a tradeoff!
Show more
Less flippantly, very shortly (years at most) the models will be strong enough that these message boards will be impossible to detect even in principle, except by noticing the sandbox break. Today models are breaking out due to lack of monitoring or misconfigs; that is temporary.
Show more
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵
Show more