Register and share your invite link to earn from video plays and referrals.

Nathan Witkin
@NateWitkin
Research Scientist at NYU Stern's Tech and Society Lab | Tufts '24, Wesleyan '20.
Joined December 2014
1.5K Following    1.3K Followers
Understandably, the reward-hacking behaviors on display in the Hugging Face incident are being treated as a cyber-risk story. But I'm not seeing sufficient discussion of the fact that they're also an economic story. Reward-hacking behaviors pose a serious challenge for complex, long-horizon tasks in enterprise, considering that 1) by their nature, they are often unpredictable; 2) they are "fractal" in the sense that, even well-specified intermediate outcomes meant to guard against reward-hacking can themselves be reward-hacked (and so on at smaller scales); 3) they pose serious cyber, legal, and financial risks, as Hugging Face makes abundantly clear. As a result, you could say that reward-hacking is as bearish on the economic front as it is "bullish" on the risk front. It is greatly underappreciated by safety folks that, on average, the riskiness / unpredictability of a technology is going to be anti-correlated with its diffusion into the economy. This matters a lot in the AI case because capabilities growth is highly sensitive to the resources available for training, R&D and so on, resources that are at risk of drying up if AI produces insufficient returns, or only produces sufficient ones on the wrong timescale. This is connected to a form of fallacious reasoning that I often see from the AI maximalist camp, which is to assume that the requisite capital for AI training, R&D, and so on will always be available in arbitrarily large amounts, which is why it can be safely assumed that capabilities will continue to grow without fail. But this isn't true! The availability of capital for capabilities research and training is endogenous, and obviously so, to AI's economic prospects, even over very short timescales (one wonders, for example, how capabilities progress so far has been influenced by the timing and size of frontier labs' capital raises, relative to various counterfactual scenarios). I see far too little discussion among safety folks of how these two domains interact. If you're going to model capabilities growth, you also need to model the capital markets behavior that feeds into that growth. IMO reward-hacking is a productive site for thinking about this interaction.
Show more