Stereo depth models usually force a tradeoff.
📌 Paper + Code
Either strong zero-shot generalization
or real-time speed.
Fast-FoundationStereo closes that gap.
The model accelerates FoundationStereo by more than 10× while keeping comparable depth quality, making real-time stereo matching practical for robotics and XR.
On an RTX 3090, it produces dense metric depth maps and 3D point clouds in real time, outperforming existing real-time stereo approaches and even challenging some heavier zero-shot models.
The key idea is a divide-and-conquer architecture that removes bottlenecks in common stereo matching components while preserving the strengths of the original teacher model.
Training also tackles a major constraint in stereo vision: the lack of ground-truth metric depth data.
The team built an automatic data curation pipeline that generates pseudo-labels from internet-scale stereo imagery using the Stereo4D dataset.
Result: a real-time stereo foundation model with strong zero-shot performance and practical runtime speed.
Code, pretrained weights, and pseudo-labeled stereo data will be released.
Thanks for sharing,
@bowenwen_me!
📍Paper:
Code:
Project:
——-
Weekly robotics and AI insights.
Subscribe free: