They do a bunch of ablations on short, long contexts and just general LM evals. QSA seems better on language modelling (again, flops matched i assume), basically the same on RULER.
Efficiency wise, its obviously better and doesn't affect MTP acceptance