Register and share your invite link to earn from video plays and referrals.

Zac Kenton
@ZacKenton1
Amplified Oversight team lead @GoogleDeepMind | AGI safety & alignment | Enabling accurate human supervision of superhuman AI.
Joined May 2014
1.6K Following    2.3K Followers
1/7 Can AI debate reduce reward hacking in RLAIF? Training against a weak LLM judge, RLAIF hacks: reward rises but judge and policy accuracy both collapse. Debate with an adversarial critic maintains judgments, allowing the policy to recover 45% of the performance gap to RLVR 🧵
Show more