RL helps models learn how to reason with different strategies, but some strategies are more effective than others.
But are the strategies learned by RL the ones that are most effective in improving accuracy?
Our new work finds that the answer is "not always"!