๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Zac Kenton
@ZacKenton1
Amplified Oversight team lead @GoogleDeepMind | AGI safety & alignment | Enabling accurate human supervision of superhuman AI.
๊ฐ€์ž… May 2014
1.6K ํŒ”๋กœ์ž‰ ์ค‘    2.3K ํŒฌ
1/7 Can AI debate reduce reward hacking in RLAIF? Training against a weak LLM judge, RLAIF hacks: reward rises but judge and policy accuracy both collapse. Debate with an adversarial critic maintains judgments, allowing the policy to recover 45% of the performance gap to RLVR ๐Ÿงต
๋” ๋ณด๊ธฐ