注册并分享邀请链接,可获得视频播放与邀请奖励。

Han Guo
@HanGuo97
PhD Student @MIT_CSAIL | Past: @togethercompute @LTIatCMU @MITIBMLab @UNCNLP, @SFResearch, @BaiduResearch | Machine Learning, NLP.
加入 August 2016
4.5K 正在关注    4.3K 粉丝
LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs).
显示更多
0
16
689
103
转发到社区