注册并分享邀请链接,可获得视频播放与邀请奖励。

Xichen Pan
@xichen_pan
PhD Student @NYU_Courant, Researcher @AIatMeta; Multimodal Generation | Prev: @MSFTResearch, @AlibabaGroup, @sjtu1896; More at
加入 August 2022
584 正在关注    746 粉丝
Modern text-to-image models are increasingly powered by large pretrained LLMs. But there is a curious mismatch: the LLM typically encodes the prompt only once, while the evolving noisy latent states are handled entirely by a newly trained generative backbone. Can pretrained multimodal prior participate in the denoising process? Introducing RepFusion. (1/12) 📄 🌐
显示更多
0
2
133
35
转发到社区