Can we train a one-step discrete generator without a teacher model?
Introducing Discrete Beckmann Transport Models 🛸 — a new family of discrete generative models that provably carries any point in the latent space to a fixed point on the vertices of the simplex in a single step.
We achieve the diversity-coherence Pareto frontier for unconditional language modeling and SOTA few-step reasoning performance, with 84.6% accuracy in 4 NFEs on Sudoku-Hard and 16.8% in 32 NFEs on TinyGSM/GSM8K.
Joint work with
@ShiyiWangML from my visit at Harvard!
Also, huge thanks to
@msalbergo and
@BrianLee794309 for the helpful discussions and guidance :D
More in 🧵