Presenting my grand unified theory of ML researcher impact: Your impact is directly proportional to how much pain you cause to infra.
Fundamentally, you can only inflict pain upon infra if your approach actually works. And the better your approach works the more pain infra is forced to endure.
So, to give some examples:
- MoE's add a ton of data-dependent computation => pain (shazeer++)
- GDN/KDA are the most complex architecture I've been forced to care about and a very annoying matrix inversion => pain (sonta++)
- Muon is much more annoying than Adam and causes annoying restrictions on parallelism => pain (keller/jeremy++)
- RL scaling forced many researchers to care about LLM inference and RL infra as a category => pain (tworek++)
Even papers like Attention Is All You Need have lead to significant pain! Before transformers were invented everyone was running small jobs and I never needed to think about kv-caches or 6D parallelism.