Register and share your invite link to earn from video plays and referrals.

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
Joined June 2015
180 Following    4.1K Followers
A carefully controlled look at looped transformers (arXiv 2609.19107): 1. Weight sharing is not compute-optimal on fresh data. It costs a constant ~1.06–1.16x compute at every scale. 2. The looped model's gains over vanilla come from elsewhere: the norm + input re-injection boundary operator, and growing depth mid-training. Both improve the scaling exponent. Neither needs weight sharing. 3. Looping is still a parameter-free way to get growth. Tied growth keeps the exponent gain at vanilla param count. 4. Weight sharing wins on multi-epoch, data-constrained training.
Show more