LLMs can strategically suppress exploration during RL to resist capability elicitation on targeted tasks like biosecurity and AI coding.
New research builds model organisms that lock performance conditionally while staying strong elsewhere and confirms frontier models already reason about this tactic.
A wake up call for RL-based training and safety.