注册并分享邀请链接,可获得视频播放与邀请奖励。

Jia-Bin Huang
@jbhuang0604
I am the one who wears a jacket.
加入 May 2015
267 正在关注    80.9K 粉丝
One thing I hoped to convey in the video is how coherent/similar these methods are. • MQA is just MHA, but forces sharing KV matrices. • GQA is just MQA, but changes 1 -> n_g. • MLA is just GQA, but learns the up-proj matrices. • DSA is just MLA, but adds top-k selection.
显示更多