Register and share your invite link to earn from video plays and referrals.

Joschka Braun
@BraunJoschka
AI safety researcher @ApolloResearch | Science of Scheming | prev. @MATSprogram @kasl_ai @health_nlp @uni_tue
641 Following    590 Followers
Only took me 8 years to hire Joschka! Among other things, our discussions about AI safety led him to switch to CS and ML back in the day :) Then during MATS he continued to do great work! Welcome to the team
Show more
1/ I recently joined @ApolloResearch in London to work on the Science of Scheming. I’m continuing my research on safety failures in LLM post-training, with my current project studying how reward-seeking develops during RL training.
Show more
I think Tübingen has become one of the best places in the world for academic AI safety research. Multiple groups across the university, MPI-IS, ELLIS Institute, and AI Center now work on AI safety & security. Having done the ML Master’s there, I can also strongly recommend it.
Show more
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
Show more
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
Show more
0
384
13.4K
2K
Forward to community
I’m giving a talk on exploration hacking with Safe AI Germany (SAIGE) today at 18:00 CEST. I’ll discuss whether LLMs can learn to resist RL training, and why this matters for post-training and capability elicitation. Join here:
Show more
Presenting today at ICML 2026 🇰🇷 Exploration Hacking: Can LLMs Learn to Resist RL Training? 10:30–12:15 KST Hall A, Poster #3101# Come by to chat about AI safety, exploration hacking, RL training dynamics, reward hacking, scheming, or capability elicitation.
Show more
I’ll be at ICML 2026 in Seoul 🇰🇷 presenting our paper: Exploration Hacking: Can LLMs Learn to Resist RL Training? Wed Jul 8, 10:30–12:15 KST Hall A #3101# Reach out if you’re interested in: - AI safety - RL training dynamics - scheming - reward hacking
Show more
RL assumes that LLMs explore well during training. What if they choose not to? In our new ICML paper with @GoogleDeepMind, we train LLMs that strategically resist RL capability elicitation by under-exploring. We study this threat model, called exploration hacking.
Show more
FPCG enables steering with comparable strength and less quality degradation, addressing a core problem with existing methods, identified by @BraunJoschka et al. in A Sober Look at Steering Vectors. FPCG also enables steering where activation steering fails completely. (7/n)
Show more
My AI Safety Paper Highlights of April 2026: - *Research sabotage propensity* - 2 sabotage benchmarks - Alignment research automation - Misaligned AI organizations - Exploration hacking - Conditional emergent misalignment More at
Show more
We looked at exploration hacking, a much-talked-about safety problem with so far ~no empirical work Bad news: we can make LLMs strongly resist RL elicitation Good news: we had to try pretty hard and it's easy to detect Excellent work led by @BraunJoschka @eyonjang @DamonFalck
Show more
Can LLMs learn to resist RL training? We empirically study exploration hacking: models controlling behavior during RL to prevent unwanted capabilities from being reinforced. Joint work with @GoogleDeepMind and @MATSprogram. More in the thread below 👇
Show more
LLMs can strategically suppress exploration during RL to resist capability elicitation on targeted tasks like biosecurity and AI coding. New research builds model organisms that lock performance conditionally while staying strong elsewhere and confirms frontier models already reason about this tactic. A wake up call for RL-based training and safety.
Show more
It was a pleasure to co-lead this joint project with @GoogleDeepMind and @MATSprogram. I found the conditionally locked model organisms on WMDP particularly interesting, and I learnt loads about RL and model organisms of misalignment. We’ll be at ICML 2026 in Seoul to present!
Show more
RL assumes that LLMs explore well during training. What if they choose not to? In our new ICML paper with @GoogleDeepMind, we train LLMs that strategically resist RL capability elicitation by under-exploring. We study this threat model, called exploration hacking.
Show more