注册并分享邀请链接,可获得视频播放与邀请奖励。

Jerry Tworek
@MillionInt
CEO and co-founder of @coreauto former VP of RL @ OpenAI : reasoning models, o3, o1, GPT4, ChatGPT, Codex, RL for robots cautious AI optimist
加入 January 2013
1.2K 正在关注    43.1K 粉丝
Interesting thing about contemporary agents is their "progressive misalignment". When they start a long running task they really try to be aligned and obey all the users intentions. They just have a tiny chance of misbehavior each step. Tiny chance of behaving out of distribution. Once they do it quickly becomes a new normal. Any tiniest bad behavior is quickly followed by more of it and it gets progressively worse the longer it takes. Reminds me of some things, but the state space of aligned behaviors seems to be unstable right now
显示更多
0
29
232
13
转发到社区