arXiv:2609.06665nvidia/nemotron-3.5-lightning-30b-a3bSeptember 6, 2026

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal EstimationExplained for Beginners

Mingwei Li, Yi Yang, Hehe Fan

Computer Vision and Pattern Recognition

Abstract

Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5° of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.

Here is the structured explanation of the research paper "TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation."

1. The Problem

To set the stage, it is helpful to understand how modern AI "sees" depth and shape. Many state-of-the-art models use a two-step process: first, they compress the input image into a compact, smaller representation called a latent (this is the job of a VAE, or Variational Autoencoder). Then, they predict geometry-related information, like surface normals (the direction a surface is facing), while working in this compressed space. Finally, they decode the prediction back into a normal map that looks right to us.

The paper identifies a specific, frustrating bottleneck in this pipeline. Because the VAE compresses the image by a factor of 8x, a lot of fine detail gets squeezed. When the model tries to predict or decode surface normals—which are crucial for understanding edges, like where a cup meets a table or where a transparent glass ends—this compression blurs the edges. The authors quantify this: even if you feed the model the "correct" answer and just ask it to encode and decode it, you still get a measurable error of 1.3 to 8.5 degrees mean angular error (MAE). At object boundaries (edges), this error balloons to 2.8 times the global average.

For a real-world user, this means AI-generated depth maps or normal maps might look smooth but will often get the "fringes" wrong—think of trying to digitally paint a transparent object or measure the angle of a blade edge, and the AI gets it subtly wrong. The paper argues that this "VAE reconstruction degradation" is an under-studied error source that limits the precision of these models.

2. How It Works (The Technical Mechanics)

TransNormal-2 is built on a framework called FLUX.2 using rectified flow. Without getting bogged down in the math of diffusion, think of rectified flow as a way to guide the AI from a chaotic mess of noise to a clean image in a single, deterministic step—like taking a short, straight path rather than a windy road. This makes the model incredibly fast (single-step inference).

The authors tackle the precision problem on two fronts: during training and at inference.

During Training: Geometry-Aware Losses Usually, models are trained to minimize the difference between their prediction and the ground truth (a simple "mean squared error"). TransNormal-2 adds three extra "losses" to guide the learning process, acting like a strict teacher:

  1. Inverse Rendering Self-Consistency: The model renders the predicted normal map back into an image and compares it to the original photo. If the rendered image doesn't match, the normal is wrong. It’s like the model checking, "Does this normal direction actually explain the lighting I see in the photo?"
  2. von Mises-Fisher Angular Loss: This is a mathematical way of saying, "The normal should point in this specific direction on a sphere." It respects the fact that normals live on a sphere's surface, punishing predictions that are close but slightly off-angle more intelligently than a flat error score.
  3. Wavelet Edge-Aware Regularization: This targets the edge problem specifically. Using wavelets (a mathematical tool for analyzing signals at different scales), the model is encouraged to be very precise at boundaries while being allowed to be slightly smoother in the flat middle of objects.

Inference: The Geometric Refinement Module (GRM) Even with great training, the "8x compression" problem rears its head at the very end when the latent prediction is decoded back into a normal map. To fix this without overturning the whole model, the authors introduce the Geometric Refinement Module (GRM).

Think of the GRM as a pair of "smart tweezers." The model gives a coarse, somewhat blurry normal prediction. The GRM looks at the original RGB image and the coarse normal map. If it sees a sharp edge in the image (like the rim of a mug) but the normal is slightly blurry there, the GRM applies a tiny, precise residual correction—nudging that specific edge pixel to be sharper. Crucially, it does this while staying grounded in the RGB data, ensuring it doesn't hallucinate normals that don't exist in the image.

3. Key Results & Benchmarks

The results are compelling, especially because they attack the "annotation scarcity" problem. Training normal estimation models usually requires thousands of images where humans have manually labeled the direction of every surface. TransNormal-2 changes the game.

  • General Scenes: On standard benchmarks, TransNormal-2 matches or exceeds the previous state-of-the-art model, MoGe-2, on all eight reported metrics. The kicker? It uses only 1.4% as many task-specific normal annotations. This means it learned to be nearly as good as the best model using a fraction of the expensive training data.
  • Transparent Objects: This is where the paper really shines. Transparent objects (like glass or plastic) are the nightmare scenario for normal estimation because you can't easily see the surface—you see through it. TransNormal-2 reduces the Mean Angular Error (MAE) by 4.2° on ClearGrasp and 3.1° on ClearPose compared to the best previous methods. That is a massive improvement for a field where a 1° or 2° gain is considered successful.

4. Why It Matters

  • Dramatic Data Efficiency: By achieving MoGe-2-level performance with only 1.4% of the normal annotations, this lowers the barrier to entry. Researchers and companies can now train high-quality normal estimators without needing massive, expensive datasets of labeled normals.
  • Transparency Breakthrough: The significant improvements on ClearGrasp and ClearPose open doors for practical applications. Better normal estimation is critical for robotic grasping (knowing exactly how to grip a slippery glass) and augmented reality (realistically rendering how light hits a glass object).
  • The "Edge" Problem is Solved: The introduction of the GRM and wavelet losses specifically targets the boundary errors that plagued previous models. For engineers, this means more reliable edge detection in autonomous driving scenarios or better surface reconstruction in 3D scanning.
  • What to Watch For: The main limitation is that the model is still bound by the physics of the VAE compression, though the GRM mitigates it. Users should watch for future work on even higher-resolution latent spaces or architectures that bypass the VAE entirely for normal estimation. The reliance on inverse rendering consistency also means the model might struggle with very unusual lighting conditions not seen during training.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →