Really cool work on data-constrained pretraining!
Key idea is to assign repeated data an effectiveness: how many fresh tokens would have produced the same validation loss? That lets them put repeated and fresh data on a common scale.
1/7
The fundamental problem with ML conference is that there is upside to doing bad work & *zero* downside. I've unearthed paper-invalidating problems with nearly every conference's Best Paper and Orals that I've touched
(two papers forthcoming on this)
1/2
People consistently underestimate how good Muse Spark is
In perceptual tasks, it beats Claude Fable in any task I've thrown at it eg
(Claude Fable incorrectly answered 6 when a friend tried)