A carefully controlled look at looped transformers (arXiv 2609.19107):
1. Weight sharing is not compute-optimal on fresh data. It costs a constant ~1.06–1.16x compute at every scale.
2. The looped model's gains over vanilla come from elsewhere: the norm + input re-injection boundary operator, and growing depth mid-training. Both improve the scaling exponent. Neither needs weight sharing.
3. Looping is still a parameter-free way to get growth. Tied growth keeps the exponent gain at vanilla param count.
4. Weight sharing wins on multi-epoch, data-constrained training.