Register and share your invite link to earn from video plays and referrals.

Ryan Greenblatt
@RyanGreenblatt
MTS at @METR_Evals, focused on using investigations to make risk-relevant information about what's going on inside AI companies public.
Joined September 2023
10 Following    22.4K Followers
We've written a post arguing that latent reasoning architectures (aka 'neuralese') would substantially increase misalignment risk via making oversight much harder. In the extreme, we could see massive 'neuralese hivemind swarms' where the agents think and communicate in latents, likely making oversight nearly entirely reliant on observing the actions these agents take. (And these agents would have huge amounts of time to reason about obfuscating their actions if they wanted to do so...) Individual agents doing extensive latent reasoning would also be concerning; in the post we discuss how above some threshold of latent reasoning, agents may be able to perform difficult-to-detect and reliable steganography for communication and further reasoning. We argue both that latent reasoning architectures would make chain-of-thought no longer very useful for oversight (by eliminating or greatly reducing the need for verbalized reasoning) and that, without these architectures, it's likely the value of chain-of-thought for oversight could be preserved.
Show more