Does 3D reconstruction have to be complex?
We answer this question with PointDiT (#
ICML2026#): a minimalist pixel-space Diffusion Transformer without bells and whistles.
We show that a plain ViT can estimate dense 3D point maps by operating directly on raw patches. No hybrid ViT+Conv architectures, no lossy VAEs, no complicated training losses. (1/5) 🧵👇