Register and share your invite link to earn from video plays and referrals.

Brendan McCord 🏛️ x 🤖
@Brendan_McCord
The academy for philosopher-builders ( A law unto myself, just like you.
3.7K Following    13.5K Followers
If AI agents were angels, no sandbox would be necessary
It’s great to see @tylercowen putting forth a concrete solution in @TheFP today, and I agree with its aims. We need frontier AI governance to be on stable footing while avoiding an FDA-style permissioning regime. If taken seriously, Tyler’s plan has a realistic shot at improving the regulatory equilibrium we’d otherwise get. I especially like his idea of using liability relief to induce better self-governance from the AI community. That said, I think the mechanism Tyler sketches out could result in an AI FDA all the same, unless a few design problems are solved. 1. Is there a market failure that justifies creating this institution? In the 1950s, uncapped liability in the burgeoning field of nuclear power was threatening private investment, so a federal liability framework called Price-Anderson was introduced to cap and pool catastrophic nuclear risk. I think Tyler needs the equivalent to be true for AI, but I don’t think that it is. My friends at frontier labs can tell me if I’m wrong, but I don’t think you are currently unable to deploy because tort exposure is uninsurable or prohibitive? If this doesn’t solve an actual bottleneck, the proposed inducement (liability safe harbor) won’t do much inducing. The institution could just sit there, irrelevant. Or, the government might decide to artificially create demand for it by making certification very valuable indeed. Make it matter, for example, for winning government contracts or for avoiding liability in court. Also, if the economic wedge isn’t compelling enough (or even if it is, I suppose), incumbents might join because controlling the standard is itself valuable. This is how a formally voluntary regime could become a de facto permissioning one. 2. What should markets do vs. what residual problem requires law? Speaking of economic wedges, I’ve gotten to see up close as an investor the way @aiunderwriting has demonstrated what a real market for part of this problem looks like. I don’t think they are the entire answer (or any one company) but I’d want that model to complement and inform the kind of thing Tyler is envisioning. For background, @RuneKvist the CEO built it after being the first commercial hire at @AnthropicAI. He saw that companies wanted to deploy agents, but their CISOs, procurement teams, lawyers, and other counterparties needed assurance regarding increasingly complex failure modes, and they needed someone to bear losses. To solve this problem, AIUC bundled technical certification with actual insurance so they can offer coverage for AI agent failures. I give this example because AIUC has something Tyler's SRO concept doesn't naturally have, which is a good (endogenous) feedback mechanism. If AIUC's standard is bad, you will see bad risk pricing on the back of it. This will mean bad loss ratios for insurers and/or high prices for customers, or you’ll have competing standards swoop in. Either way, AIUC won’t be around very long, nor should it. Its authority originates in market transactions and is sustained by them. If AIUC continues to build a great standard (and I believe they will!), the same feedback loop helps it stay fresh. If there’s an incident/near miss, that feeds directly back into underwriting. In this way, you get a standard that can update on the relevant (very fast) timescale of AI, as opposed to whatever cycle time the political version could muster. I don’t want to overstate the case for liability insurance. I don’t think it’s a “pure” signal of underlying risk. Insurance of this kind partly prices whatever the court counts as harm, which is often determined by the government. My point is that this is a _much_ higher quality feedback mechanism vs. a government-created legal privilege, which is basically what I envision Tyler’s institution offering. Nor am I suggesting that it’s obvious that an AIUC-style mechanism can price catastrophic failures from recursive self-improvement or frontier lab x-risk scenarios. But there, I would argue there is a bigger private market governance design space than ordinary commercial insurance. You could imagine labs pooling risk through a mutual insurer that provides coverage based on meeting XYZ conditions, where those entail audits and incident reporting, etc. As mentioned, that gives you a great way to have underwriting respond to technical evidence like near misses or the Huggingface incident before a court has ever decided that something like that should count as compensable harm. I would ask Tyler to pin down where private risk-pooling of this kind, or even more creative alternatives, actually falls short, so we can see where the residual problem is that requires a government-created safe harbor. As I see it, we need a tiered system in which you first let competitive underwriting and certification from companies like AIUC do the fine-grained pricing and discovery of risks the market can bear, then for risks that can’t be handled by ordinary commercial insurance you test risk-pooling with mutual insurers. For anything else that is truly residual, you use the safe harbor from Congress. I suspect that last category is not very large or urgent today, and I’m not sure it’ll ever be. And here I would want multiple qualifying regimes rather than one industry SRO, where market-based activities at the first two levels can surface competing standards, so the law can then recognize them without anointing a singular body to decide what should count. 3. Is this the least bad political equilibrium? I don’t actually think Tyler is convinced there is any market failure. I think he is making a deeper bet that we can build a relatively liberal, technically competent institution now, before a major incident gives Washington an excuse to build something much worse. I find that argument somewhat compelling (and it also makes me glad that my day job is philosophy not policy!). But I think it makes constitutional design more and not less important. It means that this institution has to be designed, anti-fragilely, to absorb pressure during a crisis. After an incident, we should expect to see government make certification mandatory or enlarge the body’s jurisdiction. The constitutional problem is how to hold the institution to its limits precisely when the political incentives to relax them are strongest. If Tyler’s wager is that building the right institution in advance gives us a better place to channel future demands for action, then the durability of those limits is a _central_ part of the mechanism, and it is missing here. 4. How do we keep certification from becoming permissioning or cartel power? I can imagine a chain of events in which you start with a totally voluntary cert, then procurement officers like it and end up requiring it, then insurers price around it, then courts treat it as evidence of reasonable care, then legislators incorporate it. You end in a situation where being “uncertified” is commercially or legally impossible and you are backdoored into an AI FDA. The same kind of accretion path can also result in a cartel. While I think accusations of regulatory capture are often overblown, I do agree with Tyler that incumbents might advocate for strict standards in part because they don't want “lower-price, lower-quality upstarts” to be able to eat their lunch. If you have a facially neutral standard that requires $20M of evals and specialized teams, you can easily imagine a burden of fixed costs that incumbents can absorb and startups cannot. And I think this will be hard to avoid, and that there will be a double-edged sword effect from some of the coordination / safety infrastructure we would definitely like to see - sharing of incident data, common threat models, etc. - also becoming a way for incumbents to increase barriers to entry. All of this makes the constitution of an institution like Tyler’s at least as important as the quality of its initial standards. I would want Tyler to say more about what powers it should get, which powers it should not get, and how it will work to challenge it. Some initial thoughts: (A) The expert body can update technical standards within an enumerated jurisdiction but doesn’t also have authority to enlarge that jurisdiction. This gives you the benefit of being able to update, say, the cybersecurity standard in response to new cybersecurity possibilities. But it blocks the ability for the body to decide that “AI safety” now also encompasses labor displacement or misinformation, or whatever else is politically salient at the time. Just because it’s a highly dynamic technology doesn’t mean we should allow a dynamically redefinable jurisdiction. (B) There should be multiple certification bodies that are able to confer the same legal benefit under transparent and objective criteria. We can’t have one incumbent-dominated SRO become the sole gateway to lawful or commercially viable deployment of AI. (C) Nonparticipation by AI labs or startups should not automatically become evidence of wrongdoing, otherwise the nominally voluntary safe harbor basically becomes a mandatory standard of care. (D) There needs to be some separation between saying “the technical standard changed” and having the legal consequences change with it. It’s very important to be able to revise technical standards rapidly, but we can’t let that become a delegated ability to revise the legal obligations attached to those standards. (E) Entrants who are contesting some standard should never have to ultimately appeal to the incumbents who wrote it; there needs to be independent contestability. And in general, no single body should simultaneously define what “AI safety” means, write the standards to implement that definition, and serve as the exclusive gateway to the legal privileges attached to complying with it. *** Tyler is right to want an expert, predictable, non-FDA-style institution, and I think his proposal is the kind of second-best institution worth trying to make work. But I'd build from competitive underwriting and certification upward and be extremely careful about creating a single federally privileged standard-setter. As I said in my response to the Pacing Letter (below), the key constitutional challenge is allowing technical standards to evolve at frontier speed without allowing the institution’s jurisdiction - or the legal power attached to its standards - to expand with them.
Show more
A lot of my friends and/or people I admire signed “Pacing the Frontier.” I think this was a bad move. My disagreement isn’t with the forecast or the framing of the coordination challenge, but with the immense and illiberal power the letter implies. There is no object called “the pace.” Progress at the frontier comes from compute, algorithms, data, post-training, inference, unattended task length, the spread of model weights, how researchers organize, and other things we haven’t invented and don’t yet know about. Inquiry leads to progress along dimensions that can’t be exhaustively specified in advance. That’s the nature of the frontier. If you gate compute, the research effort moves to algorithms. Regulate releases? Labs start taking things in-house. And other 2nd order effects will be unpredictable. Any rule that must pace the frontier involves ever-shifting proxies. It requires that its administrator has standing authority to continually redefine what counts as dangerous progress. What else is required beyond adaptive scope? The pacing regime would also need speed. One can’t successfully intervene on recursive self-improvement only after six months of legislation and litigation. It will require executive discretion. The pacing regime would also need under-the-hood access. Frontier progress is a process. The regime would need to see internal model use, training activity, compute infrastructure, and perhaps code -- proprietary and strategically sensitive information. And the thresholds couldn’t be fully public, lest they invite firms to game them. So some standards and evidence would remain secret. Insofar as the regime had to verify a rival state’s compliance, that would be an intelligence function. Restrictions would be triggered partly by evidence an affected company or researcher, or the public, could not inspect. Because this contemplated power cannot be bounded by a stable regulatory object (in the way, say, nuclear weapons can be), it would depend heavily on discretion, speed, internal access, and secret evidence. This has a highly illiberal character. Coercive power should be specific, limited, reviewable, and governed by general and knowable rules. Its characteristics (e.g., trigger, scope, evidentiary standard, duration, exceptions, means of review) should be stated before the power is granted. And the burden is on those who would propose it. A defender might answer that the proposed tool need not be coercive at all. That it could be narrow and advisory, focused on evaluation and transparency and readiness. But that wouldn’t solve the letter’s stated problem: racing. With race dynamics, each actor is under pressure not to slow down because others may continue (and thus the frontier keeps advancing). You need a mechanism to bind defectors. Voluntary norms tend to be great for binding people and firms that interact repeatedly and care about reputation. But the letter says each company and _country_… and you can’t rely on informal solutions when dealing with an unwilling state. That’s why the audience for this letter is Washington and why it calls for an international effort. Its diagnosis implies a binding mechanism. @deanwball thinks it is sensible to have a break-glass plan. That plan must involve a binding instrument, because nothing weaker addresses the problem the letter describes. But that therefore carries the burden for the use of coercive power, mentioned earlier. @johnschulman2's suggestion that labs design voluntary mechanisms among themselves is a different notion and coherent one (I would have signed that letter), but the word “country” makes this direction incompatible with the pacing letter. @OpenAI recently argued that a federal evaluator shouldn’t be able to block deployments. A week after, @AnthropicAI proposed that the government should be able to block deployments. Both labs endorsed the same letter. Whether or not the state may stop a deployment is a central question. Yet the letter accommodates both positions. What then, does the letter really say? Like the “We Must Act Now” letter from @erikbryn, @ajay_bcv, @akorinek, and @testingham before it, the letter secures agreement at an altitude where the main disagreement disappears. Lastly, the benefit of pacing is not established. The kind of slowdown the signatories have in mind would seek to buy us time for things like alignment, cyber defense, biological countermeasures, or scientific understanding -- things that increasingly depend on technologies a pause would restrict. E.g., Anthropic's framework relies in part on AI-based biological countermeasures and its security program uses AI to give defenders an advantage. A researcher in the letter's own friendly commentary was astonished at how much agents accelerated the work of the best alignment people he knows, and gave that as his reason for wanting six more months. When danger and our capacity to respond to that danger are plausibly both accelerating, the relevant question is whether this relationship is asymmetric in a safety-improving direction at the level of real-world risk. A slowdown needs to differentially slow the production of danger vs. our capacity to understand and contain that danger. The letter doesn’t attempt to establish that. It treats slower and safer as though they are the same; they are not. The letter is a serious warning, but it is no good as a warrant for an undefined power over inquiry.
Show more