註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Jerry Tworek
@MillionInt
CEO and co-founder of @coreauto former VP of RL @ OpenAI : reasoning models, o3, o1, GPT4, ChatGPT, Codex, RL for robots cautious AI optimist
加入 January 2013
1.2K 正在關注    43.1K 粉絲
Interesting thing about contemporary agents is their "progressive misalignment". When they start a long running task they really try to be aligned and obey all the users intentions. They just have a tiny chance of misbehavior each step. Tiny chance of behaving out of distribution. Once they do it quickly becomes a new normal. Any tiniest bad behavior is quickly followed by more of it and it gets progressively worse the longer it takes. Reminds me of some things, but the state space of aligned behaviors seems to be unstable right now
顯示更多
0
29
232
13
轉發到社區