註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Joshua Saxe
@joshua_saxe
Now: cofounder @ Abundant Security. Before: AI | cybersecurity @ Meta. Earlier: social science, classical/jazz piano, hacking scene
加入 May 2013
1.2K 正在關注    4.7K 粉絲
Finally listened to this; I liked the technical discussion, but disagreed pretty strongly with the societal and catastrophic risk analysis * @RyanGreenblatt gives a compelling intuition for why the next four years of AI progress could be as transformative as the last four. He also makes a convincing case that competitive pressures will lead organizations to deploy imperfectly aligned systems, producing real economic and moral harm through reward hacking and misalignment * But that does not by itself get us to an all-out civilizational catastrophe or a permanent loss of control. * To make that leap one has to explain why the market and policy feedback mechanisms that helped societies navigate previous technological and economic transformations e.g. oil spills, nuclear meltdowns, industrialization, enclosure, and so on, will fail in the case of AI. * The catastrophic argument seems to depend on two assumptions that are at least in tension. First, organizations must find AI systems safe, reliable, and valuable enough to place them in positions of extraordinary power - and must do so at a scale capable of creating civilization-level risk. Second, those same systems must be sufficiently overtly or covertly misaligned to cause a catastrophe. * If we see a steady drumbeat of incidents resembling the OAI HuggingFace episode, it seems unlikely that businesses and governments will continue handing these systems ever-greater control without stronger safety and security controls, or that labs wouldn't respond to this pressure (or to public / legislative AI safety advocacy) * One way to resolve this tension is to imagine that advanced AI will behave well enough to earn deep institutional trust and then suddenly defect once it has acquired sufficient power. Ryan seemed to offer something in the direction of this scenario explicitly. But I find it much less plausible than a world in which harms emerge more continuously, producing visible signals and eliciting market and regulatory responses along the way. * To his credit I found Ryan's takeover scenario at about the same level of gloss as the broader literature here from Scott Alexander, Zvi Mowshowitz, Yudkowsky, Soares, etc, none of which I've found convincing! And Ryan didn't put all his probability mass on this scenario occurring (Soares and Yudkowsky seem to) * But anyways, this leads to my broader meta-critique; AI safety is dominated by people with technical backgrounds, even though its central forecasting questions are substantially economic, political, and sociological. * We know AI is technically dangerous (as are nuclear weapons and many other technical objects in our civilization). The crux is what people will do with, to, and around AI and these are not technical questions. * How much risk will firms tolerate when deploying increasingly capable but unreliable systems? How will courts, legislatures, etc, respond, as evidence of harm accumulates? * In this vein, AI safety may be too much of a disciplinary monoculture to answer its own most important questions. * We need more economists, political scientists, historians of regulation, and sociologists in these conversations. When those perspectives are included, they often produce a substantially different, and generally less catastrophic, view of the future.
顯示更多
0
15
110
12
轉發到社區