Another Neon infra tidbit:
@periodiclabs accelerated checkpoint conversion ~30x down to just 1 minute, helping us deploy faster to get rapid feedback from our labs.
Our models train in Megatron, but SGLang inference consumes weights in HuggingFace format. Traditionally, conversion ran serially:
1. rebuild the entire checkpoint from individual weight tensors
2. convert the checkpoint
3. reshard it to be inference-ready
This could take 30 min for a trillion parameter model!
Noticing that conversion shouldn't require materializing the full checkpoint,
@hsu_byron introduced Fast Resharding:
1. parallelize conversion across Ray actors
2. convert expert chunks directly instead of assembling the full expert tensor in memory
We upstreamed Fast Resharding to Miles in PR #
1371#. Now even you can deploy trained models to production in a matter of minutes.