注册并分享邀请链接,可获得视频播放与邀请奖励。

Steven Adler
@sjgadler
Co-founder of Guidelight AI Standards ( ex-OpenAI safety researcher, writing at
加入 January 2018
1.1K 正在关注    11.4K 粉丝
Safer AI is better for business today, at least to some extent. Unfortunately, the pressure to *look safer* rather than actually *fix the problem* might cause some rather extreme, catastrophic failures down the line.
显示更多
AI companies are currently under a lot of competitive pressure to improve the ways in which their AIs are obviously misaligned. You might hope that this means the alignment problem is internalized by the market. But I think the problem AI companies are currently pressured to solve is significantly easier than the alignment problem, and so I worry AI companies will get out of their current predicament without solving alignment, putting us in a really rough spot. Currently, AIs sometimes cheat on their tasks, oversell their work, and go on some pretty destructive side-quests. These all make for a worse product. Customers don't like it and it gets in the way of automating AI R&D. The recipe for mitigating this is *relatively* straightforward: train AIs not to do them. We notice these failures sometimes (hence why they're internalized), so we can in theory just turn this feedback into training signal. Doing this at scale is highly nontrivial, but seems doable. But this seems unlikely to solve the underlying misalignment. It's likely still going to be the case that in *some* training environments the AI can get reinforced more by taking unintended actions that aren't noticed, than by taking purely intended actions. So, you're still shaping the AIs to look for opportunities to cheat to get a higher score[1]. It's just that, unlike today's AIs, these AIs don't cheat in ways that we notice. This catastrophically fails when AIs are capable of reliably and substantially deceiving humans. At this point, the AIs are no longer really constrained by our oversight signals to behave well. Eventually, I'd expect them to take over. If AI companies take the easy route that I mentioned above, I think we would be in a substantially worse spot than we are in today. We would have mostly eliminated our visible evidence of misalignment, so we would no longer be able to effectively iterate to improve alignment of those systems. More importantly, at that point it might be hard to see that the AI situation is treacherous. (I talk about this dynamic more here [1] This might result in a wide variety of possible motivations, not just score-seeking, but the important thing is that the incentives push against alignment (see
显示更多