Question for LLM enthusiasts
these days, one typically kicks off post training by distilling some reasoning trajectories from some other model. it's kind of like a sourdough starter
so how was the *first* model post-trained? did humans write the original set of formatting traces from scratch?
(this is why Inkling used a small number of Kimi traces btw)