One thing I hoped to convey in the video is how coherent/similar these methods are.
• MQA is just MHA, but forces sharing KV matrices.
• GQA is just MQA, but changes 1 -> n_g.
• MLA is just GQA, but learns the up-proj matrices.
• DSA is just MLA, but adds top-k selection.
顯示更多