註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Jakub Pachocki
@merettm
OpenAI
加入 August 2018
26 正在關注    93.8K 粉絲
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
顯示更多
0
274
6.4K
508
轉發到社區