註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Tri Dao
@tri_dao
Asst. Prof @PrincetonCS, Chief Scientist @togethercompute. Machine learning & systems.
加入 May 2012
661 正在關注    44K 粉絲
As hybrid models (Qwen 3.5 / Nemotron Ultra) run agents with massive context, Gated-DeltaNet / Mamba states become a bottleneck. A simple insight to make this 2x faster: load the states, compute, but don't store them. This recompute trick finally unlocks spec decoding for SSMs
顯示更多
Why do we store the SSM state at all? More and more models are hybrids (Nemotron-3, Qwen3.5), so SSM decode speed matters. We only write it back every step so the next step can read it. ReplaySSM caches the recent inputs instead and rebuilds the state on the fly. Same outputs, half the memory traffic → ~2x on spec decode at large batch sizes, which barely even helped SSMs before → up to 1.43x standard decode on large hybrids (up to Nemotron-Ultra-550B) Work with @tri_dao Blog + Code👇
顯示更多
0
5
357
39
轉發到社區