.
@arimorcos opens "Data Quality Is the Compute Multiplier" on a squeeze most teams aren't pricing in: H100 prices reversed their multi-year decline and sit roughly 40% above last year's lows, reasoning models burn eight times the tokens non-reasoning models do, and that's projected to 5x again. Google capped Meta's Gemini usage over inference constraints. OpenAI is effectively selling token futures. Ari, co-founder and CEO of DatologyAI, argues the way out isn't more compute, it's better data, and the talk is on
@aiDotEngineer's YouTube.
What you get from it is a set of measured results showing how far curation alone moves a model, plus the mechanics of how it's done.
- Curation changes the exponent, not just the intercept. The Beyond Neural Scaling Laws work showed that choosing data well bends the scaling curve itself, because you stop spending compute on redundant tokens.
- Four Cs: clean, curate, create, compose. Heuristic filters and benchmark decontamination first, then quality classifiers, redundancy reduction and task distribution matching, then synthetic data, then mixing and sequencing across training stages.
- No universal golden dataset. A dataset is only optimal with respect to the tasks you want out of it, so relevance, diversity and correct mixing are the levers.
- Vision language models. Curating a roughly 25 billion token adapter dataset moved error about 14 absolute percentage points and came within a point of a 4B Qwen model while using 145x less training compute.
- Concision turns out to be a data property. Models trained on the curated data gave much shorter responses, landing similar accuracy at around 35x fewer flops per correct answer.
- Multilingual results off 8% multilingual tokens. Most languages had at most 6 billion tokens, and curating the English data alone lifted non-English accuracy, with transfer strongest for languages most similar to English.
- Rephrasing avoids model collapse. All the information comes from the source document, so the rephraser only has to transform it, which means you can train a model better than the one doing the rephrasing. Which documents you rephrase matters more than how many.
- Mid-training multiplies post-training. Thomson Reuters saw about five points on LegalBench from a 100 billion token mid-train, under 1% of the pre-training budget, with general capabilities holding. Their unchanged post-training harness then produced close to triple the delta.
- Small runs de-risk big ones. Two trillion-token dense models drew a line running through RCAI's Trinity Large, a 17 trillion token sparse MoE trained with 50x the compute, and RCAI reached the open frontier for under $20 million all in.
I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!