check out RAEv2 led by Jas. through extensive exps, we found some really intriguing behaviors showing why strong representation encoders are key for pixel decoders.
spoiler: it’s not about hillclimbing fid; new metrics like ep
@fid-k/fdr^k show there’s a lot more left to explore!
In Oct last year, Representation Autoencoders provided an elegant solution to unified tokenization for understanding and generation.
Today we make them a bit more simple. a bit more general.
Result: >10x faster convergence, better reconstruction, better generation.
And yes we test them on T2I and world models :)
Introducing RAEv2
もっと見る