註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Haofei Xu
@haofeixu
PhD student @ETH_en and @uni_tue. Training neural networks to solve 3D, 4D, and world models.
加入 February 2022
758 正在關注    1.3K 粉絲
Does 3D reconstruction have to be complex? We answer this question with PointDiT (#ICML2026#): a minimalist pixel-space Diffusion Transformer without bells and whistles. We show that a plain ViT can estimate dense 3D point maps by operating directly on raw patches. No hybrid ViT+Conv architectures, no lossy VAEs, no complicated training losses. (1/5) 🧵👇
顯示更多
0
11
403
66
轉發到社區