登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Eyon Jang
@eyonjang
AGI safety researcher (MATS 8)
参加 December 2012
1.2K フォロー中    131 ファン
Can LLMs learn to resist RL training? We empirically study exploration hacking: models controlling behavior during RL to prevent unwanted capabilities from being reinforced. Joint work with @GoogleDeepMind and @MATSprogram. More in the thread below 👇
もっと見る