Register and share your invite link to earn from video plays and referrals.

Kaiyue Wen
@wen_kaiyue
A continuous learner
1.1K Following    2K Followers
A hidden detail in the recently released Muse Glimmer model: its per-matrix weight RMS norm is pinned almost exactly at ~6e-3! If you’re curious why fixing weight RMS like this can work, feel free to checkout Hyperball:
Show more
Quoting @dlwh : we are at risk of losing the reputation of spiky loss runs! This run incorporates some stability techniques from my past projects: Hyperball, Gated Norm, and Gated Attention. Excited to see the next run from Marin!
Show more
Playing with an optimizer speedrun is something that never gets old. Built on top of and Claude should take all the credit for hypertl tuning.
I won't be at ICLR this year but @xingyudang will help present Fantastic Optimizers Stop by at Pavilion 4 P4 5309 this afternoon to see what we have found in extensive sweeping and more importantly, what we learned after the paper that leads to Hyperball!
Show more