arXiv:2609.03756nvidia/nemotron-3.5-lightning-30b-a3bSeptember 3, 2026

ENEAS: Embedding-guided Neural Ensemble for Adaptive SegmentationExplained for Beginners

Javier del Pino, Salvador Rodríguez, Alejandro Garabito +2 more

Computer Vision and Pattern RecognitionArtificial Intelligence

Abstract

We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities. ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas

Here is the structured explanation of the ENEAS paper, written in a conversational yet precise tone.

1. The Problem

Most people interacting with AI image tools have felt the frustration of a "smart" model getting things wrong in obvious ways. The paper highlights three specific, persistent failures in today’s leading segmentation models—like SAM 3—when dealing with real-world video or image collections:

  • Temporal Hallucinations: If an object leaves the screen, the model often hallucinates it back. It might decide that a curtain or a person’s hair is actually the target object because it looks somewhat similar.
  • Spatial Fragmentation: When a camera zooms in extreme close-ups, these models tend to segment only the high-frequency textures (like brushstrokes or fabric patterns) rather than the object as a whole. They lose the "forest for the trees."
  • Semantic Misclassification: This is perhaps the most interesting failure. These models are visually oriented. If you show them a hyper-realistic statue or a large pop-art painting and ask them to segment "a person," they will often comply. They prioritize visual features—like facial symmetry or limb-like shapes—over the ontological reality that a statue is not a living person.

The authors argue these failures stem from a root cause: strong perception coupled with weak verification. A model might see a shape, but without a mechanism to say "this is not what we are looking for," it defaults to guessing based on pixels alone. This is a significant problem for applications like 3D reconstruction, where mistaking a visitor for a sculpture, or a painting for a person, literally corrupts the digital asset being built.

2. How It Works (The Technical Mechanics)

ENEAS (Embedding-guided Neural Ensemble for Adaptive Segmentation) isn't just one model; it’s a "two-way" system built on shared components. At its core, it uses Florence-2 to understand the prompt and SAM 2.1 to generate the actual masks. The magic happens in how it decides what to do with those masks.

A. Instance Tracking (The "Follow That Object" Mode)

If you give ENEAS a specific instance to track (e.g., "the blue painting"), it doesn't just look at every frame independently. It builds on a robust architecture called SeC, which previously only accepted point clicks.

ENEAS extends SeC with a "text-prompting adapter." Here is the simple flow:

  1. Initialization: On the first frame, the system uses Florence-2 to find the region matching your text prompt.
  2. Propagation: It then hands that region off to the SeC tracker. SeC has "temporal memory"—it remembers what the object is conceptually.
  3. The Trick: Because SeC has this memory, if the object disappears behind a column or leaves the frame, the tracker doesn't panic and start segmenting the closest look-alike. It simply outputs a "null" mask (True Negative), preserving the identity until the object returns. It also keeps the object "whole" even if the camera zooms in extreme close-ups, because it relies on the conceptual memory rather than just local pixel matching.

B. Semantic Discovery (The "Find All Instances" Mode)

If you give ENEAS a category prompt (e.g., "find all the chairs"), it works like a cascade of filters, moving from cheap checks to expensive checks.

  • Step 1: Visual Proposal (The Net). The system uses Florence-2 to cast a wide net, proposing every possible region in the frame that might be a chair. It is deliberately permissive; it’s better to propose a table than to miss a real chair.
  • Step 2: Embedding Verification (The Speed Check). Each proposal is scored against the text prompt using SigLIP 2, a vision-language embedding model. Think of this as a fast "rough guess." To make this robust, the system averages the score across several slightly different wordings of the prompt (e.g., "chair" vs "seat").
  • Step 3: The Uncertainty Interval. The scores are split by two thresholds. High-score proposals are accepted immediately. Low-score ones are rejected immediately. The ones in the middle—the "uncertain" ones—are passed to the next stage.
  • Step 4: Semantic Verification (The Judge). Only the ambiguous candidates go to a small Vision-Language Model (Qwen3-VL), acting as a judge. Crucially, the paper describes a clever trick here: the model is shown the candidate region, but neighboring pixels are masked out. The judge is asked: "Is this exactly a chair, and not a statue or a picture of a chair?" It must answer in a fixed binary format, which is much faster than free-form reasoning. Because this stage is only reached when the fast visual filter is unsure, the system keeps latency low.

3. Key Results & Benchmarks

The paper's most striking numbers come from the Church Statues dataset, a custom capture designed to break models by placing hyper-realistic sculptures next to real people.

  • vs. SAM 3: On the prompt "person," SAM 3 achieves a Precision of only 11.1%. This means for every real person it correctly finds, it hallucinates approximately eight artifacts (statues, paintings). Its F1 score is a low 19.5%.
  • ENEAS 2B: The proposed method jumps the F1 score to 82.8%. More importantly, it maintains a Precision above 94%. It successfully filters out the statues and paintings while recovering most of the actual people.
  • ENEAS 4B: Scaling up the verification model to 4 billion parameters pushes the F1 to 87.6% and Recall to 79.6%, nearly matching the baseline's recall while keeping precision near-perfect.

On the SA\text{}Co/VEval benchmark (a more standard video tracking test), ENEAS improves the "Association" metric (AssA) significantly over SAM 3, meaning it is much better at keeping track of the same object over time without losing it or swapping identities.

4. Why It Matters (Key Takeaways)

  • Reliable 3D Reconstruction: This is the paper's primary motivation. In 3D reconstruction (like Gaussian Splatting), you need to separate the "keep" from the "remove." If the model mistakes a visitor for a statue, that visitor becomes a permanent "ghost" in the 3D scene. ENEAS provides the semantic discrimination needed to cleanly remove people while preserving artifacts.
  • The "Two-Threshold" Trade-off: The ablation studies reveal a crucial practical insight. You can tune the system. If you are processing a standard video of people walking, you can set the thresholds for "Fast Mode," activating the expensive VLM judge rarely (achieving near real-time speeds with near-perfect accuracy). If you are processing a museum collection with tricky statues, you switch to "Robust Mode," accepting a slower per-frame time for dramatically higher semantic purity.
  • Visual Priors vs. Cognitive Reasoning: The paper effectively diagnoses the failure mode of current foundation models. They are excellent at matching shapes but poor at understanding "what things are." ENEAS bridges this by using the fast model for shape matching and the VLM for "common sense" verification only when the shape match is ambiguous.
  • Limits to Watch: The authors are upfront about boundaries. The system is only as good as its initial proposals (if the detector misses a tiny, distant object, ENEAS can't save it). Also, in extreme extreme close-ups where there is zero surrounding context, the VLM judge can be fooled just as easily as the visual model. However, these limitations are framed as "active areas" for future work rather than deal-breakers.

Summary: ENEAS moves the field from "perception-only" segmentation to "perception-plus-verification." It offers a practical, configurable way to ensure that AI models don't just see a shape, but understand what that shape actually represents, making them viable for high-stakes tasks like 3D reconstruction and precise video analysis.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →