註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
加入 June 2015
180 正在關注    4.1K 粉絲
I think we should literally stop calling every linear model as an SSM. Mamba2: Sₜ = αₜSₜ₋₁ + kₜvₜᵀ GDN: Sₜ = αₜ(I − βₜkₜkₜᵀ)Sₜ₋₁ + βₜkₜvₜᵀ GDN-2: Sₜ = (I − kₜ(bₜ⊙kₜ)ᵀ)DₜSₜ₋₁ + kₜ(wₜ⊙vₜ)ᵀ GDN family is a gradient step on a local regression loss, not a discretized ODE. Umbrella term should be "linear RNNs", with SSMs as one sub-family.
顯示更多
0
3
158
20
轉發到社區