reve 2.1 is an upgrade that we've been cooking since last month!
i had a lot of fun (and stress hehe) scaling its pretraining with async DP. it is a cross-cluster distributed algorithm that allowed us training with more compute, while staying within our existing provision. we implemented it from scratch and derisked across a few scaling rungs. it matches sync DP after getting knobs right, costs minimal overhead, and offers elasticity at the same time. cool stuffs 🦋
(also sharing a comic from our model)
顯示更多
Reve 2.1 is here.
The world’s best 4K image model just got better.
Greater prompt understanding, world knowledge, and stronger foreign-text rendering.