Everything you need to know about model distillation:
Open models like Kimi K3 and DeepSeek V4 now match frontier models on some benchmarks. The explanation everyone reaches for is distillation.
Before you buy that explanation, there are 3 questions to answer:
What is distillation, really?
What kind is even possible on a frontier model?
And can it explain the recent success of open models?
1. Traditional distillation (the real thing)
Training a student model on the full outputs of a teacher model. Not just it’s output - its reasoning traces and its confidence on every next word: what it chose, what it almost chose, and by how much.
Done right, the student inherits the teacher's performance at a fraction of the cost.
But this requires full access to the teacher's internals. Only a lab distilling its own models can do it. Nobody outside the building gets the probability distributions.
2. Cross-lab "distillation" (what people actually mean)
Frontier models stopped returning reasoning logic last year. They hide or heavily summarize the chain of thought before outputting.
What's left is behavior parroting: training on finished outputs rather than the full decision-making process.
Useful synthetic data? Yes. True distillation? Not close. You cannot do end-to-end distillation of a locked-down frontier model through an API.
3. So does distillation explain the success of open models?
No.
K3 launched weeks after Fable, and 7 days after GPT-5.6. You can't extract enough high-quality signal through a restricted API to match frontier performance - especially in that short a window.
Distillation is inevitable and it isn't a sufficient explanation.
As
@GavinSBaker put it, the "Sputnik moment" for Chinese open source hasn't arrived yet. When it does, open models will be more performant AND cheaper than closed ones.
The gains in performance are undeniable. And they can’t be attributed to distillation.