arXiv:2608.24738nvidia/nemotron-3.5-lightning-30b-a3bAugust 25, 2026

TorchMorph: CUDA-accelerated Morphological TransformsExplained for Beginners

Kai Zhao

Computer Vision and Pattern Recognition

Abstract

Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e. scipy.ndimage, is CPU-only, single-array, and therefore unusable inside a GPU training loop without an expensive device-to-host round trip. GPU vision libraries built on PyTorch cover a narrow subset of these operators, typically restricted to two spatial dimensions and flat structuring elements. We present TorchMorph, a lightweight PyTorch extension that closes this gap. TorchMorph exposes 22 public operators covering binary morphology, greyscale morphology, exact and approximate distance transforms, and entropy-regularised optimal transport, all implemented as fused CUDA kernels that operate directly on (B, C, Spatial...) CUDA tensors with up to eight spatial dimensions. The API deliberately mirrors scipy.ndimage argument-for-argument, including border modes, structuring-element origins and pre-allocated outputs, so that existing pipelines port with a change of import. We describe the layered architecture and the kernel designs behind each operator family. Against single-threaded CPU references, batched execution reaches up to 1.1e3 times the throughput of scipy.ndimage on greyscale morphology and up to 350x on exact Euclidean distance transforms, while the Sinkhorn solver runs up to 42x faster than POT. Binary and chamfer operators reproduce their SciPy counterparts exactly, and every float-valued operator agrees with the CPU reference to within 1.8e-6 absolute error. TorchMorph is released under the MIT licence at https://intcomp.github.io/tm.

The Problem

If you work with 3D medical images, satellite imagery, or any high-dimensional volumetric data inside a PyTorch training loop, you’ve likely hit a frustrating wall. The standard-bearer for morphological operations in Python is scipy.ndimage. It is excellent for offline analysis, but it was designed for a CPU world. It lives in host memory as a NumPy array. Whenever you call it from inside a GPU-accelerated pipeline, you are forced to perform an expensive round-trip: copy the data from the GPU to the CPU, run the operation, and copy it back. This serializes your training step, kills performance, and makes it impractical to use these powerful shape-processing tools where they are needed most.

Furthermore, existing GPU libraries often fall short. They might offer 2D morphology, but they typically lack the full scipy.ndimage argument set (border modes, structuring-element origins), they don't support more than two spatial dimensions, or they are restricted to flat, simple structuring elements. For researchers working with volumetric data—say, 3D segmentation masks or 4D time-series tomography—these limitations are deal-breakers. They need a tool that operates natively on tensors, supports arbitrary spatial ranks, and plays nicely with batching.

TorchMorph steps into this gap. It is a lightweight PyTorch extension that brings the breadth of classical morphological transforms into the GPU era, allowing these operations to happen inside the training loop without the costly CPU handoff.

How It Works (The Technical Mechanics)

At its core, TorchMorph is a wrapper around highly optimized CUDA kernels. The magic lies in how it bridges the gap between the familiar scipy.ndimage API and the raw hardware.

The API Mirror The most impressive engineering decision is the API design. TorchMorph argues argument-for-argument with scipy.ndimage. If you are used to calling scipy.ndimage.grey_dilation(input, size=3), you can swap the import to import torchmorph as tm and call tm.grey_dilation(input, size=3) with almost zero code changes. It handles the argument normalization—resolving structuring elements, mapping border modes, and checking origins—in a Python layer so the C++/CUDA code below doesn't have to worry about these details. This "argument-for-argument" compatibility means existing pipelines can port over with a simple import change, as demonstrated in the paper's Listing 1.

The Kernel Architecture Under the hood, the library is organized in three layers. The bottom layer is the Kernel Layer, consisting of fused CUDA kernels. The design philosophy here is to avoid "death by a thousand cuts." A naive GPU implementation would recompute coordinate mappings and boundary checks for every single output element, which is incredibly slow.

TorchMorph optimizes this heavily:

  1. Host-side Precomputation: Before the data hits the GPU, the structuring element is flattened into a list of active entries. For each entry, it precomputes the per-axis offset and a flat offset relative to the input's spatial strides. Inactive footprint positions are simply discarded before the kernel launches.
  2. Device-side Interior Path: On the GPU, each thread first checks if its output coordinate is "interior" (away from the edges). For the vast majority of voxels, this is true. These interior threads take a "fast path" that adds the precomputed flat offset directly to the linear index. There are no per-axis arithmetic operations and, crucially, no boundary test. This turns what would be a complex loop into a single addition operation.
  3. Edge Handling: Threads near the image boundaries take a slower, general path that resolves out-of-bounds neighbors to match scipy.ndimage's five border modes (constant, reflect, nearest, mirror, wrap).

This design ensures that the heavy lifting is done once on the host, while the GPU kernel is lean and fast.

Distance Transforms and Optimal Transport For distance transforms, TorchMorph uses a separable lower-envelope algorithm. Imagine computing the distance by sweeping across the image one dimension at a time, each time computing the "lower envelope" of a family of parabolas. To parallelize this, TorchMorph assigns one thread block per scanline. A single thread builds the envelope in shared memory, and the rest of the block cooperates to load the line and query the result in parallel.

For the Sinkhorn solver (entropy-regularized optimal transport), the architecture is equally clever. It tackles the problem of repeated matrix reads. Instead of re-reading the costly cost matrix for every batch item, it tiles the batch and streams the matrix row once, applying it to multiple scaling vectors held in registers. It also uses CUDA graphs to replay iteration batches after an initial compilation phase, reducing launch overhead.

Key Results & Benchmarks

The numbers speak for themselves, translating raw throughput into tangible impact.

Throughput Gains The paper reports significant speed-ups over single-threaded scipy.ndimage. On grey-level morphology, batched execution reaches up to 1,100 times the throughput of SciPy. On exact Euclidean distance transforms, the speed-up is up to 350 times. To put this in perspective: an operation that took a CPU several milliseconds per slice now takes microseconds on the GPU for a whole batch.

Numerical Faithfulness A common concern with GPU porting is precision. TorchMorph addresses this head-on:

  • Binary and Chamfer operators reproduce their scipy.ndimage counterparts exactly.
  • Float-valued operators (greyscale morphology, distance transforms) agree with the CPU reference to within 1.8e-6 absolute error. This is float32-level precision, meaning the results are numerically identical for all practical engineering purposes.

Optimal Transport Compared to POT (Python Optimal Transport), the TorchMorph Sinkhorn solver runs up to 42 times faster on grid-structured problems, making differentiable optimal transport viable for real-time training loops.

Why It Matters (Key Takeaways)

  1. End the Round-Trip: By operating directly on (B, C, Spatial...) CUDA tensors with up to eight spatial dimensions, TorchMorph eliminates the GPU-to-CPU round-trip. Morphological operations can now sit directly in the critical path of a training loop, enabling real-time shape processing.
  2. True N-D Support: Unlike many vision libraries stuck in 2D, TorchMorph handles up to eight spatial dimensions. This makes it viable for 3D volumetric data, 4D time-series, or higher-dimensional tomography without special-casing.
  3. API Familiarity: The deliberate mirroring of scipy.ndimage means the barrier to entry is low. You don't need to learn a new tensor convention or rewrite your preprocessing scripts; a change of import is often all that's required.
  4. Differentiable Transport: The inclusion of a batched, differentiable Sinkhorn solver opens the door to using optimal transport as a loss function for tasks like histogram comparison, point cloud alignment, or attention-based matching, all within PyTorch's autograd framework.

What to watch for: The authors note three current limitations. The morphology and distance kernels are forward-only (not differentiable, though subgradients exist for erosion/dilation). The modules require CUDA (falling back to SciPy is the intended path for CPU-only setups). And the computation happens in float32 with --use fast math, which introduces the tiny numerical discrepancies noted above. Future work aims to add full autograd support for morphology and expand the library of reconstruction operators.

TorchMorph represents a significant infrastructure leap, turning classical mathematical morphology from an offline afterthought into a first-class citizen of the modern deep learning pipeline.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →