RL assumes that LLMs explore well during training. What if they choose not to?
In our new ICML paper with
@GoogleDeepMind, we train LLMs that strategically resist RL capability elicitation by under-exploring.
We study this threat model, called exploration hacking.