註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Han Guo
@HanGuo97
PhD Student @MIT_CSAIL | Past: @togethercompute @LTIatCMU @MITIBMLab @UNCNLP, @SFResearch, @BaiduResearch | Machine Learning, NLP.
加入 August 2016
4.5K 正在關注    4.3K 粉絲
LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs).
顯示更多
0
16
689
103
轉發到社區