注册并分享邀请链接,可获得视频播放与邀请奖励。

Horace He
@cHHillee
@thinkymachines Formerly @PyTorch "My learning style is Horace twitter threads" - @typedfemale
加入 February 2010
612 正在关注    53.4K 粉丝
Presenting my grand unified theory of ML researcher impact: Your impact is directly proportional to how much pain you cause to infra. Fundamentally, you can only inflict pain upon infra if your approach actually works. And the better your approach works the more pain infra is forced to endure. So, to give some examples: - MoE's add a ton of data-dependent computation => pain (shazeer++) - GDN/KDA are the most complex architecture I've been forced to care about and a very annoying matrix inversion => pain (sonta++) - Muon is much more annoying than Adam and causes annoying restrictions on parallelism => pain (keller/jeremy++) - RL scaling forced many researchers to care about LLM inference and RL infra as a category => pain (tworek++) Even papers like Attention Is All You Need have lead to significant pain! Before transformers were invented everyone was running small jobs and I never needed to think about kv-caches or 6D parallelism.
显示更多
0
67
1.8K
109
转发到社区