8/ With that, we reframed multimodal generation as structured text/code generation. Diffusion just renders pixels. Planning, logic, reasoning all live in the LLM — so training looks like normal LLM training, and inherits all benefits of it: data + model scaling, reasoning, RL, tool use.
顯示更多