๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Haofei Xu
@haofeixu
PhD student @ETH_en and @uni_tue. Training neural networks to solve 3D, 4D, and world models.
๊ฐ€์ž… February 2022
758 ํŒ”๋กœ์ž‰ ์ค‘    1.3K ํŒฌ
Does 3D reconstruction have to be complex? We answer this question with PointDiT (#ICML2026#): a minimalist pixel-space Diffusion Transformer without bells and whistles. We show that a plain ViT can estimate dense 3D point maps by operating directly on raw patches. No hybrid ViT+Conv architectures, no lossy VAEs, no complicated training losses. (1/5) ๐Ÿงต๐Ÿ‘‡
๋” ๋ณด๊ธฐ