登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Bronson Schoen
@BronsonSchoen
参加 December 2024
2.1K フォロー中    969 ファン
I think Astra is not “pretty monitorable in absolute terms”. It doesn’t seem like OpenAI claims to be able to catch things like (1) sandbagging on various safety research or (2) early signs of unintended beyond episode preferences. A criticism I’ve had (at the time and now) of OpenAI’s monitorability work is that while the evals are a good starting point and the papers are clear about their serious limitations, the risk is that they’d be used to justify changes that effect monitorability and that the earlier acknowledged limitations would be ignored. From what I can tell in public posts, the reason OpenAI is confident Astra’s significantly degraded monitorability wasn’t due to architectural changes appears to be: 1. OpenAI wanted to try out new architectural changes 2. They had enough reason to suspect these changes may harm monitorability, thus ran the relevant evals 3. The monitorability evals did not surface issues that were severe enough to block the architectural changes To be clear: 1. This is *not* a claim that OpenAI has/hasn’t implemented “recurrant neuralese” (it’s unclear to me what the specific “architectural change” is) 2. There might genuinely end up being some other root cause to why the controllability differences in particular are so large However, I’m worried that we we’re making monitorability tradeoffs based on monitorability metrics we know ahead of time would be misleadingly optimistic in important cases.
もっと見る