๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Joschka Braun
@BraunJoschka
AI safety researcher @ApolloResearch | Science of Scheming | prev. @MATSprogram @kasl_ai @health_nlp @uni_tue
๊ฐ€์ž… April 2020
641 ํŒ”๋กœ์ž‰ ์ค‘    590 ํŒฌ
Iโ€™ll be at ICML 2026 in Seoul ๐Ÿ‡ฐ๐Ÿ‡ท presenting our paper: Exploration Hacking: Can LLMs Learn to Resist RL Training? Wed Jul 8, 10:30โ€“12:15 KST Hall A #3101# Reach out if youโ€™re interested in: - AI safety - RL training dynamics - scheming - reward hacking
๋” ๋ณด๊ธฐ
RL assumes that LLMs explore well during training. What if they choose not to? In our new ICML paper with @GoogleDeepMind, we train LLMs that strategically resist RL capability elicitation by under-exploring. We study this threat model, called exploration hacking.
๋” ๋ณด๊ธฐ