가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
가입 June 2015
180 팔로잉 중    4.1K 팬
A carefully controlled look at looped transformers (arXiv 2609.19107): 1. Weight sharing is not compute-optimal on fresh data. It costs a constant ~1.06–1.16x compute at every scale. 2. The looped model's gains over vanilla come from elsewhere: the norm + input re-injection boundary operator, and growing depth mid-training. Both improve the scaling exponent. Neither needs weight sharing. 3. Looping is still a parameter-free way to get growth. Tied growth keeps the exponent gain at vanilla param count. 4. Weight sharing wins on multi-epoch, data-constrained training.
더 보기