Register and share your invite link to earn from video plays and referrals.

Sara Dragutinovic
@sara_drag
ML PhD @ NYU | Maths & CS @ Oxford
Joined May 2021
225 Following    208 Followers
Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization.
Show more