Turns out everyone's been initializing LoRA weights suboptimally.
We find that by using the singular value decomposition (SVD) of the weight matrix to determine both the frozen weights and LoRA initialization, we can achieve faster convergence in RL training.
See the full post below on our extension of PiSSA
At Trajectory, we're constantly implementing and building upon the latest research ideas on the path to continual learning.
We wish we had the time to share all of them, but here's a quick glimpse on our explorations with PiSSA, and choosing the right trainable geometries for agentic RL.