Register and share your invite link to earn from video plays and referrals.

wh
@nrehiew_
eng primarily, ml mostly, research previously
Joined October 2023
104 Following    18.5K Followers
They do a bunch of ablations on short, long contexts and just general LM evals. QSA seems better on language modelling (again, flops matched i assume), basically the same on RULER. Efficiency wise, its obviously better and doesn't affect MTP acceptance
Show more