๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Sebastian Raschka
@rasbt
ML/AI research engineer. Ex stats professor. Author of "Build a Large Language Model From Scratch" ( & reasoning (
๊ฐ€์ž… October 2012
1.2K ํŒ”๋กœ์ž‰ ์ค‘    511K ํŒฌ
Nice case study on using optimized functions whenever possible (except for educational purposes, though ๐Ÿ˜†)
By switching from a hand-rolled GELU to PyTorch's built-in one, I improved my LLM training speed -- and much more than I expected, from 21,000 tokens/second to 25,000!