arXiv:2609.08084nvidia/nemotron-3.5-lightning-30b-a3bSeptember 8, 2026

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth EstimationExplained for Beginners

Igor Pavlovic, Thiemo Wandel, Anton Obukhov +6 more

Computer Vision and Pattern RecognitionMachine Learning

Abstract

Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. In this paper, we revisit Marigold, a set of techniques for repurposing modern image generation and editing models, powered by the diffusion transformer (DiT) architecture, into state-of-the-art monocular depth estimators. Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run. We analyze the artifacts of naive training and identify two effective remedies: aligning the model's internal representations with semantic features extracted from ground-truth, and adopting a 2-stage fine-tuning protocol built around a novel Sinkhorn-based loss. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D. Qualitatively, our model resolves fur, foliage, and hair-thin edges that have eluded prior models. Furthermore, Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition. Project website: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web

The Problem: Why Another Depth Estimator?

Monocular depth estimation is the art of guessing the 3D structure of a scene from a single 2D photograph. In the real world, this is incredibly difficult because a flat image loses almost all information about distance. If you see a building in a photo, you don't know if it's a massive skyscraper far away or a small model sitting on your desk.

For years, this has been a bottleneck for technology we use every day. Autonomous cars need to "see" the road ahead to avoid obstacles; augmented reality apps need to know where to place a virtual object so it looks like it's actually sitting on your coffee table; and robotics relies on depth to navigate cluttered environments. Despite the field being mature, current state-of-the-art models often struggle when faced with "out-of-distribution" data—think of a self-driving car encountering a type of foliage or a texture it was never trained on. Furthermore, many models produce depth maps that are blurry or lack the sharp edges needed to distinguish a thin branch from the sky behind it.

The core gap this paper addresses is the tension between "generalization" and "detail." Most models are either great at handling unseen scenarios but produce blurry, inaccurate maps, or they are sharp but brittle, failing the moment the input changes slightly. The authors argue that we can have both, but we need to rethink how we train the models.

How It Works: The Marigold V2 Engine

Marigold V2 isn't a brand-new model architecture from scratch; it’s an evolution of a technique called "Diffusion Transformers" (DiTs). If you are familiar with AI image generation (like Stable Diffusion), you know these models are usually trained to remove noise from images over many steps. Marigold V2 takes a pretrained diffusion model and repurposes it to predict depth.

Here is the technical mechanics broken down:

1. The DiT Architecture Instead of using older Convolutional Neural Networks (CNNs) that look at pixels in small blocks, Marigold V2 uses a Transformer. Think of a Transformer as a system that looks at the entire image at once and understands how different parts relate to each other. This "global awareness" is what allows the model to generalize better to new types of scenes.

2. Single-Step Inference (The "Cheap to Run" Part) Standard diffusion models are slow because they require dozens of steps to refine an answer. Marigold V2 introduces recipes to perform "single-step inference." Imagine trying to guess the answer to a complex math problem. Usually, you’d write down an answer, check it, erase it, and try again 50 times. Marigold V2 is like having a genius who can look at the problem once and write the correct answer immediately. This makes the model incredibly fast and cheap to run, which is essential for real-world apps like robotics where latency matters.

3. The Two Remedies: Fixing the "Blurriness" The authors identified two specific problems with naive training and came up with clever fixes:

  • Aligning Representations: The model was losing track of "semantic" information (what objects are) while focusing on pixel-level details. The fix was to force the model to align its internal features with semantic features extracted from the ground-truth depth maps. It’s like telling the model, "Don't just look at the pixels; remember that this fluffy part is actually a tree."
  • The Sinkhorn-Based Loss (The 2-Stage Fine-Tuning): This is the technical heart of the improvement. The authors introduced a two-step training process. In the first stage, the model learns broadly. In the second stage, they use a "Sinkhorn-based loss." Without getting too deep into the math, Sinkhorn is a method used in optimal transport—think of it as a way to efficiently match things into groups while minimizing cost. In this context, it helps the model better align predicted depth boundaries with actual object edges. It’s essentially a sophisticated "quality control" step that ensures the depth map edges are crisp and correct.

Key Results & Benchmarks: The Numbers

The proof is in the numbers. On the KITTI dataset (a standard benchmark for autonomous driving), Marigold V2 showed a 16–26% improvement in AbsRel over the previous best model.

What does AbsRel mean? AbsRel (Absolute Relative error) is a metric that measures the average difference between the predicted depth and the actual depth, relative to the actual depth. If a model has an AbsRel of 0.1, it means, on average, its depth predictions are off by 10%.

Translating that to plain language: Marigold V2 is answering 16% to 26% more questions correctly on the standard test used by the field. That is a massive jump in a domain where improvements are usually measured in fractions of a percent.

Qualitatively, the paper highlights that the model resolves "fur, foliage, and hair-thin edges" that have eluded prior models. If you’ve ever tried to get a depth map to correctly separate a person’s hair from a blurry background, you’ll appreciate this achievement.

Why It Matters: Key Takeaways

  • Faster and Cheaper: By enabling single-step inference, Marigold V2 makes high-quality depth estimation viable for real-time applications on edge devices (like smartphones or robots) without needing massive computing power.
  • Versatility: The authors demonstrate that the model isn't just a "one-trick pony." When applied to other tasks—like estimating surface normals (the direction a surface faces) and intrinsic image decomposition (separating an image into lighting, shading, and albedo/color)—it achieves state-of-the-art results. This suggests the underlying DiT methodology is robust across different visual tasks.
  • The "Hair" Problem is Solved: For practitioners, the most immediate takeaway is the qualitative improvement in fine details. This means better AR experiences, more accurate 3D reconstructions of complex scenes, and better obstacle avoidance for drones.
  • Watch this space: The introduction of the Sinkhorn-based loss is a methodological contribution. Researchers will likely look at this loss function as a new tool for training models to have sharper boundaries.

Summary

Marigold V2 takes the powerful, flexible architecture of Diffusion Transformers and applies it to the challenging problem of seeing depth from a single image. By optimizing for speed (single-step inference) and precision (Sinkhorn loss and semantic alignment), it pushes the state-of-the-art forward by leaps and bounds. It isn't just a marginal improvement; it resolves long-standing issues with fine detail like hair and foliage while making the technology fast enough for real-world deployment.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →