注册并分享邀请链接,可获得视频播放与邀请奖励。

Tianwei Yin
@TianweiY
Multimodal AGI. Prev: @MIT @Adobe
加入 January 2019
156 正在关注    1.4K 粉丝
8/ With that, we reframed multimodal generation as structured text/code generation. Diffusion just renders pixels. Planning, logic, reasoning all live in the LLM — so training looks like normal LLM training, and inherits all benefits of it: data + model scaling, reasoning, RL, tool use.
显示更多