๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Xuanchi Ren
@xuanchi13
Ex-Senior Research Scientist @NVIDIAAI. PhD @UofTCompSci. Working on GenAI and world models
๊ฐ€์ž… March 2021
590 ํŒ”๋กœ์ž‰ ์ค‘    1.5K ํŒฌ
The latent-vs-pixel debate misses the point. GPT Image 2 shows what users notice: pixel-level fidelity. Latent models show what scales: compact semantic structure. We connect them by replacing VAE/RAE decoders with a Pixel Diffusion Decoder. Code and Model available: ๐Ÿงต(1/N)
๋” ๋ณด๊ธฐ