가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Neal Khosla
@nealkhosla
I judge myself on what my 8 and 80 year old selves would think. CEO of @CuraiHQ - making healthcare available and affordable to everyone
가입 February 2010
1K 팔로잉 중    36.4K 팬
sakana is incredible
For over a decade, we’ve accepted that end-to-end backprop is the only way to train deep networks. But holding the entire network in memory all at once is why AI training is hitting a resource wall. We found a new way to break the network into blocks and train them independently. The trick? Treating the network’s forward pass like a diffusion model denoising a signal. This reinterpretation slashes the memory needed to train deep models. In our #ICLR2026# paper ( we matched end-to-end performance across ViTs, DiTs, and LLMs. We did this while training just one isolated block at a time.
더 보기