reve 2.1 is an upgrade that we've been cooking since last month!
i had a lot of fun (and stress hehe) scaling its pretraining with async DP. it is a cross-cluster distributed algorithm that allowed us training with more compute, while staying within our existing provision. we implemented it from scratch and derisked across a few scaling rungs. it matches sync DP after getting knobs right, costs minimal overhead, and offers elasticity at the same time. cool stuffs 🦋
(also sharing a comic from our model)