註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Jia-Bin Huang
@jbhuang0604
I am the one who wears a jacket.
加入 May 2015
267 正在關注    80.9K 粉絲
One thing I hoped to convey in the video is how coherent/similar these methods are. • MQA is just MHA, but forces sharing KV matrices. • GQA is just MQA, but changes 1 -> n_g. • MLA is just GQA, but learns the up-proj matrices. • DSA is just MLA, but adds top-k selection.
顯示更多