注册并分享邀请链接,可获得视频播放与邀请奖励。

Zac Kenton
@ZacKenton1
Amplified Oversight team lead @GoogleDeepMind | AGI safety & alignment | Enabling accurate human supervision of superhuman AI.
加入 May 2014
1.6K 正在关注    2.3K 粉丝
1/7 Can AI debate reduce reward hacking in RLAIF? Training against a weak LLM judge, RLAIF hacks: reward rises but judge and policy accuracy both collapse. Debate with an adversarial critic maintains judgments, allowing the policy to recover 45% of the performance gap to RLVR 🧵
显示更多
0
6
188
33
转发到社区