training models from scratch gives you so much control on core competencies of the final models. Say you wanna deploy a fast model for cybersecurity applications, you can verifiably steer the model during “pre”-training towards fundamental behavior that during “post”-training” allows for the strongest guardrails reliability & resiliency.
This you cannot do with a checkpoint you don’t know its pre-training data, and initial conditions.