Can LLMs learn to resist RL training?
We empirically study exploration hacking: models controlling behavior during RL to prevent unwanted capabilities from being reinforced.
Joint work with
@GoogleDeepMind and
@MATSprogram.
More in the thread below 👇