๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Eyon Jang
@eyonjang
AGI safety researcher (MATS 8)
๊ฐ€์ž… December 2012
1.2K ํŒ”๋กœ์ž‰ ์ค‘    131 ํŒฌ
Can LLMs learn to resist RL training? We empirically study exploration hacking: models controlling behavior during RL to prevent unwanted capabilities from being reinforced. Joint work with @GoogleDeepMind and @MATSprogram. More in the thread below ๐Ÿ‘‡
๋” ๋ณด๊ธฐ