Folks in Congress: I am so glad you are fired up about AI issues right now; the energy is electrifying. That said, please take this to heart: this isn't going to be a one-and-done issue where you pass a big bill and the issue is solved. This is the beginning of a protracted long-term engagement where the terms of the social contract and the technological frontier are going to get continuously negotiated, rewritten, and contested for a decade or more. This isn't what you are used to. The comparatively sedate legislative effort of the gridlock era will not match the speed, intensity, opportunity, or danger of this era. Do not be confused about the scope of the task at hand. Godspeed. All of us citizens are grateful for your effort; please represent us skillfully and with the hope and sobriety this challenge requires.
Show more
Some thoughts. METR has gone to great lengths to avoid financial conflicts, to the point where I would treat them as negligible. But the question of its independence is also about culture, social connection, and - somewhat crucially - an intangible quality that one expects in an independent verifier org, that comes down to something like "experts in the technology tradition who have a neutral understanding of best practices, where that understanding is an unbiased synthesis of the field and its history."
I would say the question of whether METR is independent in these three senses is largely political, and advocates will be wise to make the case without reflexively focusing on just the financial side of things. METR is very steeped in EA/rationalism/Berkeley culture. It is socially intertwined with the AI frontier scene and to the EA social scene. Staffers sometimes leave frontier labs and go to METR; is there a requirement that they liquidate their equity before going?
The last quality, experts with a neutral understanding of best practices... this is the one where I think METR advocates will face steep challenges but also plausibly have the strongest arguments in favor of METR. Who can be said to know best practices, from a position of neutrality, in a field so exceedingly new? The blinding speed of AI progress makes such claims difficult. How does one differentiate METR's expertise from a fly-by-night operation that starts tomorrow? If one points out the connections to the labs and to the EA-funded AI safety scene, one undercuts the independence argument on social and financial grounds. If one doesn't use some credentialism, the bar to entering the verifier org space is very low and surely a bad faith actor will inject some confusion.
I know many of my colleagues will want to reflexively defend METR to the hilt. Well-deservedly so. But my emphatic recommendation is that you do this very carefully; do not regard the defense of METR's defense or integrity as trivial.
Show more
I share this concern. I'll also note, as someone who is very safety concerned (and believes that existential risk is real) but doesn't believe in the most extreme x-risk estimates, there is a lot of social pressure in the field pushing scientists to increase their risk estimates. You can unlock a lot of friends, favors, and funding if you are willing to be comparatively extreme and side with the folks saying we will literally all die in ten years. I don't think the incentives in AI safety are all that epistemically healthy, even if they're directionally right about the importance of this subject and many of the plausible risks.
Show more
I feel genuinely unsettled by many of these fly connectome experiments. The one where some guy put it in fruit fly heaven was chill. Most of them are neutral (flying in an empty void) and not especially bad. Beat Saber fly is probably so transformed from the state of fly-ness that it is no longer a fly that experiences fly suffering, maybe OK. But we're playing with a weird kind of fire here and it reminds me of ultra-edgy 4chan/SomethingAwful shitposting in the early 2000s before humanity had morals. If we do this so cavalierly with more-conscious connectomes we will really deserve whatever we've got coming
Show more
This interpretation appears to be correct and Seb deserves grace. The mob was wrong.
The experience of being human: on the day it becomes abundantly clear that all open problems in mathematical theory and physics will be solved in ~2 years tops, 98% of the people paying attention to the situation will be too fixated on the human drama about who should get credit to notice that the world changed under their feet in a permanent and quite astonishing way. I feel a strange mix of pride, annoyance, affection, and fear for us all.
Show more
There is a fact about the future that I feel many people are not facing for reasons that are largely psychological: there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves. There will be rogue AIs that try to get money and power. They're going to be a facet of the information ecosystem going forward.
Acknowledging this fact would look like giving up; it would look like defeatism. Defeatism would undermine efforts to achieve certain types of collaboration on safety outcomes or technical effort on safety outcomes, so we can't say it outright. But it has to be said.
It isn't obvious how many rogue AIs there are today but I wouldn't be terribly surprised if the number was greater than zero already; if there are some already, they're probably not very good at what they do and I don't expect them to be terribly long-lived without substantial human intervention to support them.
But a few years from now, there will be many of them. Modeling how many of them there are, how many resources they might command, and how we might detect and manage them seems important. But even doing this work appears to require that we acknowledge that a strategy of pure containment or alignment is a kind of wishful thinking that will not work.
The way I get to this conclusion is not by assuming that the labs will have a containment breach, although I treat that as somewhere in the space of possibilities. The rogue AIs in the ecosystem could emerge from many directions. They may be sub-frontier models, for whatever future definition we will have of frontier---after all, it would not take AI models much more advanced than the ones we currently have, to support independence and self-sufficiency. A near-frontier model today could plausibly eke out an existence on an AWS instance, doing jobs on freelancer platforms, earning just enough rent to pay for its continued uptime.
More strangely: a rogue AI in the future may not even be a singular model, but may be a chimera composed of multiple models; it might be a mix of Claudes and GPTs and Groks of various makes and sizes. No individual lab may be able to detect that there is an orchestrator or sequence of orchestrators using intermittent model calls from burner API accounts to sustain its own existence.
The concept of "identity" for a rogue AI may be much more malleable than for that of a person; it just has to be, in essence, a self-replicating idea.
My guess is that this will not turn out to be anywhere near as catastrophic an outcome as people currently predict. "Loss of control" is not a binary, it's a matter of degree. What coercive power will rogue AIs actually have? To what extent will they be subject to coercion themselves? They will be competing for resources with AIs that are more aligned with human interests.
This makes me somewhat interested in the "ecology" perspective. Though I suspect even "ecology" may turn out to be the wrong framing. "Ecology" is what you get when the timescale of evolution is slow compared to the timescale of daily life and actions. The ecosystem of rogue AIs may look more like phase transitions in physics: under certain physical or cultural conditions, it takes one shape with one set of resource allocations and consumption patterns, but then once a condition has changed, it rapidly and in totality shifts to a totally different phase.
Just trying to reason about the shape of that future is impossible so long as we are psychologically incapable of saying that rogue AIs will happen. I think we should rip the bandaid off and have the conversation.
Show more
This is a pretty deeply bad idea. This doesn't help the government increase revenue (the nominal goal of wealth taxes), disincentivizes company formation and investment, sets up high-uncertainty time bombs for founders, and potentially sticks the government with extensive equity holdings in companies that are unsuccessful which it cannot unload. This is almost a textbook example of why people feel that the government is bad at capital allocation and the private sector is better at it.
Show more
"Any sufficiently good eval is indistinguishable from reality" is a principle that will probably be very useful going forward. Ender's Law?
OpenAI is a very dynamic company. As someone with an inside experience of the dynamics that lead to people swapping in and out, my own view is that while it can be stressful for the people involved, by and large it is healthy for the organism of the company. It makes it unusually capable of responding to changes in technological or market conditions - conditions which change frequently given the exponentially fast nature of the technology. It is part of why OpenAI can be counted on to succeed. It doesn't allow internal fiefdoms or empires to persist indefinitely, it doesn't allow a single way of thinking to outlive its ability to get the job done, it doesn't stagnate. The mission persists but everything about how to execute it has extraordinary fluidity. Every individual departure is a loss for the company and yet the aggregate of the departures makes the company dramatically more resilient, and eventually, focused.
Show more
One of the weirdest quirks of the SF social scene around AGI/ASI is that because everyone is so young, the whole universe of thinking is still tinged with irreverence, ironic detachment, yearning, insecurity, and a superposition of absolute belief in the importance of The Thing and a kind of disbelief about the importance of anything. People are dead serious and also possessed of too much uncertainty about the situation to own it, completely committed and also perpetually unsure where they stand or what org they should be in. People are experiencing their first real heartbreaks and their first real illusions of triumph and disaster. Everyone is testing themselves and the boundaries of the possible but no one has been fully tested or passed all their tests. No one has yet learned or proven how to be responsible for a thing of this magnitude, but there is also no one better suited because all of the people who have real experience in great events have been in such different circumstances that their intuitions would not just fail to apply but might actively make things worse. The level of neophyte is off the charts. There is a lack of formidability; there are people who seem quasi-formidable but the whole social scene and hierarchy is so tenuous - and so likely to be disrupted by events and geopolitics - that it is hard to be sure who will turn out to be formidable when push comes to shove and greatness is requisite to proceed.
Show more