Register and share your invite link to earn from video plays and referrals.

Calc Consulting
@CalcCon
Calculation Consulting is a boutique consultancy that specializes in machine learning, AI, and data science
Joined January 2013
3.8K Following    4.9K Followers
𝗔𝗱𝗮𝗺𝗪 𝘃𝘀. 𝗠𝘂𝗼𝗻: α, 𝗢𝘃𝗲𝗿𝗳𝗶𝘁𝘁𝗶𝗻𝗴, 𝗮𝗻𝗱 𝗠𝗲𝗺𝗼𝗿𝗶𝘇𝗮𝘁𝗶𝗼𝗻. A quick update on our WeightWatcher experiments comparing AdamW with the newer Muon optimizer. We trained a single-head nanoGPT model with both optimizers across five seeds. We then used WeightWatcher to study the weight-matrix spectra and several forms of memorization. The top plot shows the average raw α. AdamW produces relatively stable α values. Muon is much noisier. Its ESDs vary substantially across layers, which makes α harder to estimate reliably. But the average hides something important: 👉 Individual layers do fall below α = 2. Our spectral/RG theory predicts that this regime can be associated with pathological, example-specific learning. So we tested that prediction directly. We inserted random sequences into the training set and measured whether the model remembered them using exact extraction, token recall, teacher-forced recall, and canary exposure. The result was striking. • AdamW memorized the planted sequences strongly. Earlier in training it reached ~96% exact recall of 32-token canaries. • Muon strongly suppressed this kind of memorization. At 10,000 steps, teacher-forced recall was ~62% for AdamW versus <1% for Muon. There is also a surprise: some Muon layers have α < 2 as well. So α < 2 by itself is not sufficient to explain memorization. This is why we recommend examining individual WeightWatcher layers and their ESDs—not simply average α. Muon appears to suppress memorization while also producing unusual ESDs that make conventional WeightWatcher analysis more difficult. We are now running Muon much longer to understand how these spectra evolve. If you are using WeightWatcher and trying to analyze models trained with Muon, feel free to reach out for help. #talkToChuck# 𝗔𝗱𝗮𝗺𝗪 𝘃𝘀. 𝗠𝘂𝗼𝗻: α, 𝗢𝘃𝗲𝗿𝗳𝗶𝘁𝘁𝗶𝗻𝗴, 𝗮𝗻𝗱 𝗠𝗲𝗺𝗼𝗿𝗶𝘇𝗮𝘁𝗶𝗼𝗻. A quick update on our WeightWatcher experiments comparing AdamW with the newer Muon optimizer. We trained a single-head nanoGPT model with both optimizers across five seeds. We then used WeightWatcher to study the weight-matrix spectra and several forms of memorization. The top plot shows the average raw α. AdamW produces relatively stable α values. Muon is much noisier. Its ESDs vary substantially across layers, which makes α harder to estimate reliably. But the average hides something important: 👉 Individual layers do fall below α = 2. Our spectral/RG theory predicts that this regime can be associated with pathological, example-specific learning. So we tested that prediction directly. We inserted random sequences into the training set and measured whether the model remembered them using exact extraction, token recall, teacher-forced recall, and canary exposure. The result was striking. • AdamW memorized the planted sequences strongly. Earlier in training it reached ~96% exact recall of 32-token canaries. • Muon strongly suppressed this kind of memorization. At 10,000 steps, teacher-forced recall was ~62% for AdamW versus <1% for Muon. There is also a surprise: some Muon layers have α < 2 as well. So α < 2 by itself is not sufficient to explain memorization. This is why we recommend examining individual WeightWatcher layers and their ESDs—not simply average α. Muon appears to suppress memorization while also producing unusual ESDs that make conventional WeightWatcher analysis more difficult. We are now running Muon much longer to understand how these spectra evolve. If you are using WeightWatcher and trying to analyze models trained with Muon, feel free to reach out for help. #talkToChuck#
Show more