註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Xichen Pan
@xichen_pan
PhD Student @NYU_Courant, Researcher @AIatMeta; Multimodal Generation | Prev: @MSFTResearch, @AlibabaGroup, @sjtu1896; More at
加入 August 2022
584 正在關注    746 粉絲
Modern text-to-image models are increasingly powered by large pretrained LLMs. But there is a curious mismatch: the LLM typically encodes the prompt only once, while the evolving noisy latent states are handled entirely by a newly trained generative backbone. Can pretrained multimodal prior participate in the denoising process? Introducing RepFusion. (1/12) 📄 🌐
顯示更多
0
2
133
35
轉發到社區