登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Jerry Tworek
@MillionInt
CEO and co-founder of @coreauto former VP of RL @ OpenAI : reasoning models, o3, o1, GPT4, ChatGPT, Codex, RL for robots cautious AI optimist
参加 January 2013
1.2K フォロー中    43.1K ファン
Interesting thing about contemporary agents is their "progressive misalignment". When they start a long running task they really try to be aligned and obey all the users intentions. They just have a tiny chance of misbehavior each step. Tiny chance of behaving out of distribution. Once they do it quickly becomes a new normal. Any tiniest bad behavior is quickly followed by more of it and it gets progressively worse the longer it takes. Reminds me of some things, but the state space of aligned behaviors seems to be unstable right now
もっと見る