登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
参加 June 2015
180 フォロー中    4.1K ファン
A carefully controlled look at looped transformers (arXiv 2609.19107): 1. Weight sharing is not compute-optimal on fresh data. It costs a constant ~1.06–1.16x compute at every scale. 2. The looped model's gains over vanilla come from elsewhere: the norm + input re-injection boundary operator, and growing depth mid-training. Both improve the scaling exponent. Neither needs weight sharing. 3. Looping is still a parameter-free way to get growth. Tied growth keeps the exponent gain at vanilla param count. 4. Weight sharing wins on multi-epoch, data-constrained training.
もっと見る