登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Jack Lindsey
@Jack_W_Lindsey
Neuroscience of AI brains @AnthropicAI. Previously neuroscience of real brains @cu_neurotheory.
参加 January 2019
265 フォロー中    19.3K ファン
I think it's unlikely that white-box techniques will be able to provide as much monitorability as CoT currently provides within a year. Currently, they are capable of catching some unverbalized thoughts / plans, but nowhere near as reliably as CoT seems to. I'm optimistic about "in a few years," since the rate of progress here is quite fast (our go-to activation decoding techniques for production monitoring were both published just in the last few months!). But even that is uncertain. Also, the mere existence of working techniques doesn't imply that they'd be widely applied, esp. if they are expensive or difficult to develop. I think we need defense in depth, and it'd be irresponsible not to work on developing + stress-testing white-box techniques, but it'd also be irresponsible to allow CoT monitorability to degrade substantially until we are confident we have a roughly-as-good working alternative. (To be clear, this isn't a dunk on OpenAI -- I have no evidence that OpenAI has allowed CoT monitorability to degrade substantially! Based on Jakub's comment, it sounds like probably not at this time). Aside: I think "mechanistic interpretability" isn't the right mental model for our current white-box monitoring approaches. It's really more like "mind reading" -- even when we can read the model's "thoughts," we're typically clueless about the underlying mechanism.
もっと見る