We looked at exploration hacking, a much-talked-about safety problem with so far ~no empirical work
Bad news: we can make LLMs strongly resist RL elicitation
Good news: we had to try pretty hard and it's easy to detect
Excellent work led by
@BraunJoschka @eyonjang @DamonFalck