Register and share your invite link to earn from video plays and referrals.

Joschka Braun
@BraunJoschka
AI safety researcher @ApolloResearch | Science of Scheming | prev. @MATSprogram @kasl_ai @health_nlp @uni_tue
Joined April 2020
641 Following    590 Followers
I’ll be at ICML 2026 in Seoul 🇰🇷 presenting our paper: Exploration Hacking: Can LLMs Learn to Resist RL Training? Wed Jul 8, 10:30–12:15 KST Hall A #3101# Reach out if you’re interested in: - AI safety - RL training dynamics - scheming - reward hacking
Show more
RL assumes that LLMs explore well during training. What if they choose not to? In our new ICML paper with @GoogleDeepMind, we train LLMs that strategically resist RL capability elicitation by under-exploring. We study this threat model, called exploration hacking.
Show more