Marigold V2 turns an image-editing DiT into a single-step model for sharp, detailed dense prediction.๐ Apache 2.0.
๐ค
๐
๐ Best zero-shot results across all evaluated depth datasets among models trained on comparable data. AbsRel improves by 16%โ26% over the previous best on KITTI and ETH3D.
๐ Fur, foliage, fine wires, and object boundaries stay crisp. The two-stage iREPA and SinkLoss recipe tackles the smoothing and flying-pixel artifacts common in diffusion-based depth models.
๐งฉ The same framework reaches SOTA results in depth completion, see-through depth, surface normals, and intrinsic image decomposition. Depth completion records the lowest RMSE across all four reported benchmarks.
โก A pretrained Qwen-Image-Edit DiT becomes a single-step dense predictor through lightweight adaptation. Training takes less than a week on one 32 GB GPU.