Register and share your invite link to earn from video plays and referrals.

Jia-Bin Huang
@jbhuang0604
I am the one who wears a jacket.
Joined May 2015
267 Following    80.9K Followers
One thing I hoped to convey in the video is how coherent/similar these methods are. • MQA is just MHA, but forces sharing KV matrices. • GQA is just MQA, but changes 1 -> n_g. • MLA is just GQA, but learns the up-proj matrices. • DSA is just MLA, but adds top-k selection.
Show more