註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Horace He
@cHHillee
@thinkymachines Formerly @PyTorch "My learning style is Horace twitter threads" - @typedfemale
加入 February 2010
612 正在關注    53.4K 粉絲
I agree with the quoted poster, but I do think the situation is a bit of the bell curve/midwit tweet. The original paper didn't really understand memory-bound kernels, and had some very incorrect assertions on performance. On the other hand, once you start really pushing performance (like matmul epilogues), the additional state tracking needed for the mean shift makes it somewhat more painful to fuse. So yes, layernorm is not 2x more expensive than rmsnorm. But it's not "free" either.
顯示更多
every research team needs to spend some time learning how their modeling code lowers down to gpu kernels example: rms norm vs layernorm industry assumes rms norm is cheaper. you don't need the x̂, it looks simpler, so it must be faster, so it got adopted but that's not true. both rms and layernorm are memory bound kernels (the amount of time it takes to get the data to gpu cores is longer than the amount of time it takes to do the computation on said cores) both take the same amount of time e2e so can probably train the model right with either, but maybe it makes a difference (e.g why diffusion models still use norms with an affine shift e.g adaLN) i get why. as open-source models get better and, inevitably, commoditized across the inference providers, overall serving speed determines user experience determines which model gets adopted but there's no reason to superstitiously avoid free things
顯示更多
0
7
208
10
轉發到社區