They also use Muon throughout the network
- Router AdamW for stability
- GR projections AdamW (suspected due to their weird shape)
- ngram table Adam without weight decay
- per head orthogornalization
They trained the model with TP so for the parameter update, they come up with a parameter assigner to balance flops across ranks. The ranks then exchange via all to all