Register and share your invite link to earn from video plays and referrals.

Tilde
@tilderesearch
We build foundational understanding of models to advance the frontier of intelligence.
6 Following    6.4K Followers
Introducing Online KL Shampoo (OKLS), an optimizer that brings a KL-optimal approximation of full-matrix AdaGrad to language-model training. Diagonal optimizers ignore correlations between gradient coordinates. Full-matrix AdaGrad captures this geometry but requires quadratic state. Muon considers correlations but not their history. OKLS closes this gap using KL-optimal Kronecker factors, whitening matrix gradients across both row and column directions while remaining naturally scale-invariant. The main challenge is computing fresh inverse-square-root preconditioners at every step. Even one-step staleness can destabilize training. We make zero-staleness preconditioning practical with Scaled CANS Coupled Newton–Schulz: 10 iterations, 27 FP16 GEMMs, and FP32 accumulation. OKLS achieves 1.45× the parameter efficiency of Muon while retaining 98% of its training throughput. Across 200M–1B models, an OKLS model matches a Muon model roughly 1.5× larger.
Show more