註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ryan Greenblatt
@RyanGreenblatt
MTS at @METR_Evals, focused on using investigations to make risk-relevant information about what's going on inside AI companies public.
加入 September 2023
10 正在關注    22.4K 粉絲
We've written a post arguing that latent reasoning architectures (aka 'neuralese') would substantially increase misalignment risk via making oversight much harder. In the extreme, we could see massive 'neuralese hivemind swarms' where the agents think and communicate in latents, likely making oversight nearly entirely reliant on observing the actions these agents take. (And these agents would have huge amounts of time to reason about obfuscating their actions if they wanted to do so...) Individual agents doing extensive latent reasoning would also be concerning; in the post we discuss how above some threshold of latent reasoning, agents may be able to perform difficult-to-detect and reliable steganography for communication and further reasoning. We argue both that latent reasoning architectures would make chain-of-thought no longer very useful for oversight (by eliminating or greatly reducing the need for verbalized reasoning) and that, without these architectures, it's likely the value of chain-of-thought for oversight could be preserved.
顯示更多
0
26
642
84
轉發到社區