Marigold V2 turns an image-editing DiT into a single-step model for sharp, detailed dense prediction.đ Apache 2.0.
đ¤
đ
đ Best zero-shot results across all evaluated depth datasets among models trained on comparable data. AbsRel improves by 16%â26% over the previous best on KITTI and ETH3D.
đ Fur, foliage, fine wires, and object boundaries stay crisp. The two-stage iREPA and SinkLoss recipe tackles the smoothing and flying-pixel artifacts common in diffusion-based depth models.
đ§Š The same framework reaches SOTA results in depth completion, see-through depth, surface normals, and intrinsic image decomposition. Depth completion records the lowest RMSE across all four reported benchmarks.
⥠A pretrained Qwen-Image-Edit DiT becomes a single-step dense predictor through lightweight adaptation. Training takes less than a week on one 32 GB GPU.