// Normalized Low-Rank Adaptation (NoRA) //
They propose a one-line change to LoRA that costs nothing and improves convergence, stability and forgetting.
LoRA initializes the up-projection to zero, which means early optimization is governed almost entirely by the down-projection. This observation tells you where to regularize.
NoRA normalizes the down-projection matrices during training. The authors also show the same normalization applied once at initialization improves standard LoRA without repeating it through training, which is the cheaper of the two options.
The benefits hold across pretraining, supervised fine-tuning and reinforcement learning. Faster convergence, better final performance, more stable training, and less catastrophic forgetting.
It adds no trainable parameters and no inference-time computation, which is what makes it broadly applicable rather than another specialized LoRA variant.
Paper: