가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Bronson Schoen
@BronsonSchoen
가입 December 2024
2.1K 팔로잉 중    969 팬
I think Astra is not “pretty monitorable in absolute terms”. It doesn’t seem like OpenAI claims to be able to catch things like (1) sandbagging on various safety research or (2) early signs of unintended beyond episode preferences. A criticism I’ve had (at the time and now) of OpenAI’s monitorability work is that while the evals are a good starting point and the papers are clear about their serious limitations, the risk is that they’d be used to justify changes that effect monitorability and that the earlier acknowledged limitations would be ignored. From what I can tell in public posts, the reason OpenAI is confident Astra’s significantly degraded monitorability wasn’t due to architectural changes appears to be: 1. OpenAI wanted to try out new architectural changes 2. They had enough reason to suspect these changes may harm monitorability, thus ran the relevant evals 3. The monitorability evals did not surface issues that were severe enough to block the architectural changes To be clear: 1. This is *not* a claim that OpenAI has/hasn’t implemented “recurrant neuralese” (it’s unclear to me what the specific “architectural change” is) 2. There might genuinely end up being some other root cause to why the controllability differences in particular are so large However, I’m worried that we we’re making monitorability tradeoffs based on monitorability metrics we know ahead of time would be misleadingly optimistic in important cases.
더 보기