All of this is powered by one unified architecture: a multimodal autoregressive diffusion transformer, pretrained from scratch.
This foundation blends the best of modern LLMs and video models, benefiting from the architectural, algorithmic, and systems advances from both areas.
顯示更多