๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Xichen Pan
@xichen_pan
PhD Student @NYU_Courant, Researcher @AIatMeta; Multimodal Generation | Prev: @MSFTResearch, @AlibabaGroup, @sjtu1896; More at
๊ฐ€์ž… August 2022
584 ํŒ”๋กœ์ž‰ ์ค‘    746 ํŒฌ
Modern text-to-image models are increasingly powered by large pretrained LLMs. But there is a curious mismatch: the LLM typically encodes the prompt only once, while the evolving noisy latent states are handled entirely by a newly trained generative backbone. Can pretrained multimodal prior participate in the denoising process? Introducing RepFusion. (1/12) ๐Ÿ“„ ๐ŸŒ
๋” ๋ณด๊ธฐ