가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Nathan Witkin
@NateWitkin
Research Scientist at NYU Stern's Tech and Society Lab | Tufts '24, Wesleyan '20.
가입 December 2014
1.5K 팔로잉 중    1.3K 팬
The essay below from Aaronson is just unmoored from reality, sorry. It is just not true, if you try to approach the question in a data-oriented fashion, rather than via vibes and speculative fiction and anecdotes, that capability is "dramatically ramping up ... every month." See attached graphs of progress over the previous ~year on several difficult knowledge work benchmarks. Also keep in mind that: 1) these are pass@1 scores, when what you should really be looking at to extrapolate to real world performance (esp. w/r/t full automation, "drop-in workers," etc.) is pass^n, i.e. how often, on average, models succeed at tasks n times in a row, rather than just once; 2) even strong benchmarks have low external validity because they do not and, in many cases, cannot model key constitutive properties of real world knowledge work. These include: field-specific regulatory and security constraints, ambiguous or conflicting requests from multiple stakeholders, uncodified context (related, for instance, to subtle expectations around how outputs are formatted, phrased, or otherwise presented), required use of rare or legacy software or other systems (about which there is little public data), idiosyncratic constraints on cost and timing, and sudden unexpected changes in any of the above. As I wish went without saying, any sub-tasks involved in causing catastrophic harm to humans would face many more and more difficult constraints than these. Even you if put all this to the side and just eyeball graphs, you will not see dramatic monthly ramp-ups in capability. You will certainly not see performance befitting the "machine God" that Aaronson claims is already arriving (??). As is often the case, the key error here is not to have appreciated extreme jaggedness. Solving Navier-Stokes is crazy impressive. It is also happening at the same time as models struggle to even approach human performance on many (not all!) of the complex tasks that white-collar professionals in finance, law, academia, and so on perform everyday (yes, I know NS was solved by an unreleased model; I promise that model will also underperform humans on these tasks). I can't stop you from turning your brain off and going "well obviously every form of knowledge work will fall soon, if Navier-Stokes did." But the balance of evidence suggests that this is a bad inference (even just within the field of math, by the way, where models still struggle, and will continue to struggle on many open problems). I'll begin to take doomers more seriously when they adopt a norm of trying to articulate why they expect models to best humans in fields much less hospitable to RLVR, where data is much scarcer, where tacit and context-specific knowledge plays a much larger role, and where iterative contact with the real world is essential. I'll add that it doesn't help that their concerns seem universally to pass through ill-formed concepts like RSI / AGI / ASI that they are myopically pattern-matching to a reality far too messy to accommodate them. This seemed like a point a lot of folks on here appreciated and even agreed with (esp. w/r/t AGI and ASI) not two months ago, and yet it seems to have gone out the window post-Coxon. Now these concepts have been re-drafted to serve as ineliminable premises in arguments to the effect that while models are not ready to kill us all yet they will soon pass the threshold of [key three letter premise], and so we should all be very afraid. I am not afraid and do not think you should be either. That is not because I deny AI is progressing at a rapid pace. It is instead because I see many signs that it will be both slower and more jagged than many expect, and because I see few if any signs of that progress outdoing the capacity of our institutions to adapt to it. I also find it frankly bizarre to begin worrying about extinction-level harm from AI when it can't yet hold a candle to the human cost of cars or drugs or guns or almost any other generically important technology. I see no evidence that the threat from AI is going to leapfrog past these and suddenly kill millions or billions. As a result, to organize and communicate around extinction-level threats at this stage strikes me as obviously counterproductive. All it will do and is already doing is delegitimizing the cause of AI safety (a cause I support), and creating all manner of rhetorical and strategic openings for its opponents. That is what happens when you throw millions of dollars at saying silly things you can't substantiate about a hot-button issue in public. If AI safety folks want to be taken seriously, it would help for them to be serious.
더 보기
Scott Aaronson writing tonight about rats, Eliezer Yudkowsky, the Road to Damascus, and the Singularity - which he now says has already started.