登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Zac Kenton
@ZacKenton1
Amplified Oversight team lead @GoogleDeepMind | AGI safety & alignment | Enabling accurate human supervision of superhuman AI.
参加 May 2014
1.6K フォロー中    2.3K ファン
1/7 Can AI debate reduce reward hacking in RLAIF? Training against a weak LLM judge, RLAIF hacks: reward rises but judge and policy accuracy both collapse. Debate with an adversarial critic maintains judgments, allowing the policy to recover 45% of the performance gap to RLVR 🧵
もっと見る