Register and share your invite link to earn from video plays and referrals.

Ryan Greenblatt
@RyanGreenblatt
Chief scientist at Redwood Research (@redwood_ai), focused on technical AI safety research to reduce risks from rogue AIs
Joined September 2023
10 Following    20.3K Followers
I agree with almost all of this and think the point being made here is important. (I'm probably less optimistic about having good mind reading techniques in a few years and I tend to think that AIs might get harder to mind read over time which adds an additional difficulty.)
Show more
I think it's unlikely that white-box techniques will be able to provide as much monitorability as CoT currently provides within a year. Currently, they are capable of catching some unverbalized thoughts / plans, but nowhere near as reliably as CoT seems to. I'm optimistic about "in a few years," since the rate of progress here is quite fast (our go-to activation decoding techniques for production monitoring were both published just in the last few months!). But even that is uncertain. Also, the mere existence of working techniques doesn't imply that they'd be widely applied, esp. if they are expensive or difficult to develop. I think we need defense in depth, and it'd be irresponsible not to work on developing + stress-testing white-box techniques, but it'd also be irresponsible to allow CoT monitorability to degrade substantially until we are confident we have a roughly-as-good working alternative. (To be clear, this isn't a dunk on OpenAI -- I have no evidence that OpenAI has allowed CoT monitorability to degrade substantially! Based on Jakub's comment, it sounds like probably not at this time). Aside: I think "mechanistic interpretability" isn't the right mental model for our current white-box monitoring approaches. It's really more like "mind reading" -- even when we can read the model's "thoughts," we're typically clueless about the underlying mechanism.
Show more