註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Sebastian Raschka
@rasbt
ML/AI research engineer. Ex stats professor. Author of "Build a Large Language Model From Scratch" ( & reasoning (
加入 October 2012
1.2K 正在關注    511K 粉絲
Nice case study on using optimized functions whenever possible (except for educational purposes, though 😆)
By switching from a hand-rolled GELU to PyTorch's built-in one, I improved my LLM training speed -- and much more than I expected, from 21,000 tokens/second to 25,000!