註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Zhifei Li
@andylizf
Incoming PhD @PrincetonCS. Building AI systems & infra. Prev @BerkeleySky & @ruc1937
加入 July 2022
116 正在關注    238 粉絲
one of the most annoying things in MLSys: tiny ops keep forcing data to move around CODA: do them before the tile leaves the chip
LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs).
顯示更多