註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Sara Dragutinovic
@sara_drag
ML PhD @ NYU | Maths & CS @ Oxford
加入 May 2021
225 正在關注    208 粉絲
Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization.
顯示更多
0
14
370
44
轉發到社區