註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Zac Kenton
@ZacKenton1
Amplified Oversight team lead @GoogleDeepMind | AGI safety & alignment | Enabling accurate human supervision of superhuman AI.
加入 May 2014
1.6K 正在關注    2.3K 粉絲
1/7 Can AI debate reduce reward hacking in RLAIF? Training against a weak LLM judge, RLAIF hacks: reward rises but judge and policy accuracy both collapse. Debate with an adversarial critic maintains judgments, allowing the policy to recover 45% of the performance gap to RLVR 🧵
顯示更多
0
6
188
33
轉發到社區