Leaving image generation to diffusion but reasoning to an LLM is ๐คฉ
Now take a step further and think about it in the context of robotics with spatial-action models.
Work combining modalities is still in the early days.
8/ With that, we reframed multimodal generation as structured text/code generation. Diffusion just renders pixels. Planning, logic, reasoning all live in the LLM โ so training looks like normal LLM training, and inherits all benefits of it: data + model scaling, reasoning, RL, tool use.