One open question for anyone finetuning models is - how "post-trainable" are different open source models?
We can theorize that models that reside in shallower loss curves are more amenable to post-training - which makes sense, given that it requires less gradient updates to modify behavior.
Then, this paper proposes a pre-training algorithm that makes a model more post-trainable.
The idea is straightforward once you wrap you head around it:
1. Find the worst policy locally around the policy pi (found by taking the inverse gradient w.r.t pi), call that pi_prime
2. Take the gradient at pi_prime (call that grad[pi_prime] )
3. Apply grad[pi_prime] to pi.
The intuition is that we're actually moving in the direction that benefits the worst model around you, meaning a model maintains post-trainability because we maintain a shallow loss landscape.
Another way to imagine this, is it effectively avoids steep pot-holes during pretraining that would lock your model in a distribution that it can't post-train its way out of.
Really great work from
@IshaanWatts18,
@CatherineL11638,
@goyalsachin007,
@jacspringer,
@AdtRaghunathan. It's our favorite paper of the month at Trajectory!