注册并分享邀请链接,可获得视频播放与邀请奖励。

Micah Carroll
@MicahCarroll
RSI Preparedness lead @openai Prev @berkeley_ai /w @ancadianadragan & Stuart Russell
加入 August 2011
820 正在关注    10K 粉丝
GPT6 is a very significant jump in capabilities, but also an important decrease in monitorability – especially under adversarial evaluation. We give many details about this in the system card. In my opinion, monitorability and control will likely become a major bottleneck for responsible AI development quite soon, given that risks from a fixed amount of residual misalignment grows together with capabilities. Aligning on acceptable monitorability outcomes and practices which robustly bound misalignment risks is a non-trivial problem, and there may be limited time to resolve uncertainty around this before we need to act. It should ultimately be in everyone’s interest to agree on shared bounds for monitorability in order to avoid races to the bottom – nobody wants extremely capable models whose alignment properties we don’t understand, and that are reliably able to cause severe real-world harm without being detected.
显示更多
0
47
740
89
转发到社区