가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Zhifei Li
@andylizf
Incoming PhD @PrincetonCS. Building AI systems & infra. Prev @BerkeleySky & @ruc1937
가입 July 2022
116 팔로잉 중    238 팬
one of the most annoying things in MLSys: tiny ops keep forcing data to move around CODA: do them before the tile leaves the chip
LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs).
더 보기