Nice case study on using optimized functions whenever possible (except for educational purposes, though 😆)
By switching from a hand-rolled GELU to PyTorch's built-in one, I improved my LLM training speed -- and much more than I expected, from 21,000 tokens/second to 25,000!