Register and share your invite link to earn from video plays and referrals.

Zhifei Li
@andylizf
Incoming PhD @PrincetonCS. Building AI systems & infra. Prev @BerkeleySky & @ruc1937
Joined July 2022
116 Following    238 Followers
one of the most annoying things in MLSys: tiny ops keep forcing data to move around CODA: do them before the tile leaves the chip
LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs).
Show more