And now for the second problem - scaling to large models. Here, we applied some tried and tested ideas:
1. PP-local and EP-local gather: Avoid redundant gather across PP groups, avoid gathering EP layers to avoid OOMs
2. Pipelined execution: Pipeline all-gather on the trainer, the weight transfer and the post-process on the inference side