註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Micah Carroll
@MicahCarroll
RSI Preparedness lead @openai Prev @berkeley_ai /w @ancadianadragan & Stuart Russell
加入 August 2011
820 正在關注    10K 粉絲
GPT6 is a very significant jump in capabilities, but also an important decrease in monitorability – especially under adversarial evaluation. We give many details about this in the system card. In my opinion, monitorability and control will likely become a major bottleneck for responsible AI development quite soon, given that risks from a fixed amount of residual misalignment grows together with capabilities. Aligning on acceptable monitorability outcomes and practices which robustly bound misalignment risks is a non-trivial problem, and there may be limited time to resolve uncertainty around this before we need to act. It should ultimately be in everyone’s interest to agree on shared bounds for monitorability in order to avoid races to the bottom – nobody wants extremely capable models whose alignment properties we don’t understand, and that are reliably able to cause severe real-world harm without being detected.
顯示更多
0
47
740
89
轉發到社區