註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

hardmaru
@hardmaru
Co-Founder and CEO @SakanaAILabs 🎏
加入 November 2014
1.9K 正在關注    436K 粉絲
For over a decade, we’ve accepted that end-to-end backprop is the only way to train deep networks. But holding the entire network in memory all at once is why AI training is hitting a resource wall. We found a new way to break the network into blocks and train them independently. The trick? Treating the network’s forward pass like a diffusion model denoising a signal. This reinterpretation slashes the memory needed to train deep models. In our #ICLR2026# paper ( we matched end-to-end performance across ViTs, DiTs, and LLMs. We did this while training just one isolated block at a time.
顯示更多
0
153
5.7K
623
轉發到社區