註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
加入 June 2015
180 正在關注    4.1K 粉絲
A carefully controlled look at looped transformers (arXiv 2609.19107): 1. Weight sharing is not compute-optimal on fresh data. It costs a constant ~1.06–1.16x compute at every scale. 2. The looped model's gains over vanilla come from elsewhere: the norm + input re-injection boundary operator, and growing depth mid-training. Both improve the scaling exponent. Neither needs weight sharing. 3. Looping is still a parameter-free way to get growth. Tied growth keeps the exponent gain at vanilla param count. 4. Weight sharing wins on multi-epoch, data-constrained training.
顯示更多
0
7
129
17
轉發到社區