注册并分享邀请链接,可获得视频播放与邀请奖励。

Haofei Xu
@haofeixu
PhD student @ETH_en and @uni_tue. Training neural networks to solve 3D, 4D, and world models.
加入 February 2022
758 正在关注    1.3K 粉丝
Does 3D reconstruction have to be complex? We answer this question with PointDiT (#ICML2026#): a minimalist pixel-space Diffusion Transformer without bells and whistles. We show that a plain ViT can estimate dense 3D point maps by operating directly on raw patches. No hybrid ViT+Conv architectures, no lossy VAEs, no complicated training losses. (1/5) 🧵👇
显示更多
0
11
403
66
转发到社区