註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

wh
@nrehiew_
eng primarily, ml mostly, research previously
加入 October 2023
104 正在關注    18.5K 粉絲
They do a bunch of ablations on short, long contexts and just general LM evals. QSA seems better on language modelling (again, flops matched i assume), basically the same on RULER. Efficiency wise, its obviously better and doesn't affect MTP acceptance
顯示更多