arXiv:2609.11929nvidia/nemotron-3.5-lightning-30b-a3bSeptember 10, 2026

SenseNova-U1.5: Towards Native Unified Visual IntelligenceExplained for Beginners

Haiwen Diao, Jiahao Wang, Chenjing Ding +62 more

Computer Vision and Pattern Recognition

Abstract

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.

SenseNova-U1.5: When the Model Both Sees and Draws

There’s a quiet tension in how today’s AI handles images. Most systems treat "looking" and "creating" as two separate acts. A vision encoder—trained on millions of pictures to spot cats, cars, or text—feeds its abstract understanding into a generator, often a diffusion model that works in a compressed latent space. The encoder compresses the image into a handful of semantic "tokens," and the generator expands them back into pixels. The two operate in different mathematical worlds, and stitching them together costs performance, introduces seams, and limits the kinds of tasks the system can handle.

SenseNova-U1.5, a new 8-billion-parameter multimodal model from researchers at OpenSenseNova, attacks this problem head-on. It proposes a "native unified" architecture: one model that sees, reasons, and generates directly from pixels, without encoders or variational autoencoders (VAEs) in the loop. The result is a system that can look at a high-resolution image, have a sophisticated conversation about it, and then generate or edit that image in the same representational space. The implications for anyone building visual AI—product managers, developers, designers—are significant.

The Problem: The Encoder–VAE Gap

To understand why SenseNova-U1.5 matters, it helps to understand the architectural divorce it’s trying to dissolve.

In the current state of the art, a typical multimodal pipeline might look like this: a CLIP-style vision encoder processes an image, converting it into a sequence of dense vectors that capture semantics but discard fine-grained pixel detail. A separate VAE decoder then takes a compressed latent representation and "hallucinates" the pixels back. This two-step process works, but it introduces friction. The encoder is optimized for classification and retrieval; the VAE for pixel fidelity. When you ask the system to do something that sits between those two—like precisely editing a specific region while preserving the rest, or rendering bilingual text pixel-perfectly within a complex layout—the separation becomes a bottleneck. The model has to translate back and forth between semantic space and pixel space, and each translation loses information.

SenseNova-U1.5 was born from the earlier SenseNova-U1, which first demonstrated that an encoder-free, VAE-free approach was feasible. The new version takes that insight and scales it. The core problem it addresses is this: how can a model perceive and create within a single, coherent representational space? The authors argue that the current modular approach—encoder plus VAE—is an unnecessary architectural compromise. By learning directly from raw pixels, the model can maintain spatial coherence from the first token to the last generated pixel.

How It Works: From Independent Pixels to Spatially Joint Reconstruction

The technical heart of SenseNova-U1.5 is a redesign of the "visual interface"—the part of the model that converts images into tokens and back again.

In the original SenseNova-U1, each visual token was decoded independently. The model predicted the RGB values for one 32×32 patch of pixels, then moved to the next patch, completely independently. At low resolutions, this was tolerable. But as the authors note, at high resolutions, the cracks show. You get visible seams between patches, texture discontinuities, and a general "grid-like" feel to the output. It’s as if you took a photograph, cut it into a grid of squares, and asked someone to paint each square blind, without seeing the neighboring squares. The result might be individually reasonable, but the overall image would feel disjointed.

SenseNova-U1.5 changes the decoding strategy entirely. Instead of independent patch prediction, it employs a "spatially coupled decoder." Here is the analogy that makes this concrete for any engineer or product mind: imagine you’re compositing a 4K video. Instead of rendering each frame independently and hoping the joins are invisible, you process the entire frame as a continuous canvas. You use spatial convolutions—essentially, blurring and sharpening operations that look at a pixel and its neighbors—to make sure the color at one edge matches the color at the adjacent edge. In SenseNova-U1.5, the model first reshapes its internal hidden states into a 2D grid, then applies a series of Pixel Shuffle upsampling stages (with 3×3 convolutions in between) to progressively reconstruct the full-resolution image. Crucially, the 3×3 convolutions allow information to flow between neighboring tokens before the pixels are finalized. Color, texture, and geometry are resolved jointly across patch boundaries.

The effect is substantial. The model still compresses every 32×32 block of pixels into a single visual token—a high compression ratio that keeps the sequence length tractable for the transformer backbone. But now, those tokens are not isolated islands. They are nodes in a spatial network, and the decoder’s convolutions ensure that the transition from one token to the next is smooth. The authors report that this reduces boundary artifacts and enhances stability, especially at their target resolution of 4K.

There is also a clever conditioning mechanism for resolution. Because the effective noise magnitude the model must denoise varies with image size, SenseNova-U1.5 explicitly embeds a "resolution-dependent noise-scale" into its diffusion process. If you’re generating a 4096×4096 image, the model knows it’s working with a higher-noise regime than a 512×512 image, and conditions its denoising accordingly. This allows the model to natively handle resolutions up to 4K without retraining from scratch.

The architectural unification doesn't stop at the visual interface. The model employs a "Native Mixture-of-Transformers" (MoT) design. In plain terms, this is a single Transformer backbone that handles both language and vision, rather than two separate networks glued together. Text tokens attend only to preceding text (causal language modeling), while image tokens attend bidirectionally to capture spatial dependencies. This asymmetric attention pattern lets the model "think" in language while "seeing" in pixels, with a shared hidden space where semantic ideas and visual evidence can interact directly. The authors emphasize that understanding and generation retain separate attention projections and feedforward modules—so the model isn't forced to use the same computational "gear" for perception and synthesis—but the shared attention layers serve as a bridge, enabling cross-stream communication without entangling the distinct optimization objectives.

Key Results: What the Numbers Say

The paper evaluates SenseNova-U1.5 across a sweeping range of benchmarks, and the results paint a picture of a model that is more than the sum of its parts.

Image Generation and Fidelity

On general generation benchmarks like GenEval and GenEval2, SenseNova-U1.5 sets a new state of the art among open-source models. In GenEval, which tests attributes like counting objects, positioning them correctly, and binding colors to shapes, U1.5 scores 0.92 overall—outperforming models that are significantly larger (like the 20B-parameter Qwen-Image) and besting its immediate predecessor, SenseNova-U1. The improvements are most striking in "compositional consistency": the model’s ability to satisfy multiple visual constraints within a single prompt. If you ask it to "place a red ball to the left of a blue cube and a green triangle in the foreground," U1.5 is notably more likely to get all three elements right in the right positions and with the right colors.

On GenEval2, which pushes harder on complex reasoning and instruction following, U1.5 again leads open-source, with particularly strong gains in attribute, counting, and verb-related generation. The paper notes that the model is "more capable of faithfully translating structured, relation-intensive prompts into coherent visual compositions, especially when multiple constraints need to be satisfied simultaneously." This is the kind of improvement that matters for real-world use: generating a diagram from a complex textual specification, or creating a marketing visual from a detailed brief.

The text-rendering story is perhaps even more compelling. On CVTG-2K, a benchmark focused on multi-region layouts and long-form bilingual text generation, SenseNova-U1.5 achieves an average score of 0.948—the best among all evaluated models. More importantly, its word accuracy remains above 0.95 even in the challenging four- and five-region settings. Previous models might nail the first line of text but fail as the density increases; U1.5 maintains fidelity across the whole layout. On LongText-Bench, it achieves leading open-source performance on both English and Chinese tracks, with particular strength in Chinese text rendering—a known pain point for many multimodal models.

Image Editing

The editing benchmarks tell a story of improved precision and balance. On ImgEdit, which covers a wide array of edit types, SenseNova-U1.5 achieves leading overall performance with substantial gains over SenseNova-U1. The improvements are consistent across categories, indicating stronger "edit execution" while maintaining source-image preservation and visual quality. On GEdit-Bench, which focuses on specific edit operations like adding, replacing, or deleting objects, U1.5 again tops the open-source leaderboard. Its strongest advantage lies in semantic consistency—the model can more reliably identify the intended edit, modify the correct visual content, and preserve the surrounding context.

On WeEdit, which benchmarks instruction adherence, text clarity, and background preservation, SenseNova-U1.5 substantially outperforms other open-source models and remains competitive with leading proprietary systems. The paper highlights its "strong instruction adherence, text clarity, and background preservation," demonstrating robust text-centric editing where requested textual changes are accurately executed while unrelated visual content remains intact.

On OmniRef-Bench, which tests multi-reference editing (using multiple source images as guidance), SenseNova-U1.5 achieves leading open-source performance. The gains are particularly pronounced in style consistency and background preservation, indicating the model can effectively exploit reference information while minimizing unintended modifications. Strong subject and pose consistency show that identity, appearance, and structural cues are reliably preserved across diverse editing conditions.

Reasoning and Interleaved Generation

Perhaps the most subtle but important result is how well the model's generative capabilities enhance its reasoning. On RISEBench, which evaluates edits requiring the model to infer temporal, causal, spatial, or logical consequences before modifying the image, SenseNova-U1.5 outperforms open-source baselines without chain-of-thought prompting, and explicit reasoning provides a further substantial improvement. The gains are most pronounced for causal, logical, and temporal edits, indicating that the model has internalized a form of "visual common sense" that it can deploy when needed.

On VBVR-Pro-Bench, which tests visual reasoning across in-domain and out-of-domain settings, SenseNova-U1.5 achieves new state-of-the-art results across cognitive faculties. Notably, it generalizes strongly to out-of-domain tasks, outperforming leading proprietary models. The paper reports that SenseNova-U1.5 achieves leading performance on RealUnify-GEU (a benchmark of generation-enhanced understanding) with a considerable improvement over the previous SenseNova model, while remaining competitive on Uni-MMMU-GaU. This suggests that the model's ability to generate images is not just an output modality, but can function as an intermediate representation that supports multimodal reasoning and understanding.

On OpenING, a benchmark for open-ended interleaved image-text generation, SenseNova-U1.5 with CoT (chain-of-thought) achieves leading overall performance, improving over SenseNova-U1 and outperforming compared proprietary pipelines. Strong results in image-text coherency, human alignment, and multi-step consistency demonstrate reliable open-ended interleaved generation.

A Note on Scaling

It is worth noting the parameter count: SenseNova-U1.5 has 8 billion parameters. In a field where bigger often seems better, this is a deliberate choice. The authors emphasize that the model's strength comes not from scale alone, but from the efficiency of the native unified architecture. By eliminating the encoder and VAE, and by designing a spatially coherent decoder, they achieve strong performance with a relatively compact model. The Configurations table in the paper lists the model as having 8.2B parameters, with a Mixture-of-Transformers architecture and a 32×32 patch size. The training compute was significant—spanning multiple stages with varying learning rates and sequence lengths—but the final model is efficient enough to be practical.

Why It Matters: Key Takeaways

SenseNova-U1.5 is not just a incremental improvement; it signals a shift in how we might build future multimodal systems. Here are four takeaways that matter beyond the academy:

  1. Native unification reduces failure modes. By sharing a single visual representation for both understanding and generation, SenseNova-U1.5 eliminates the "translation loss" that comes from passing between encoder features and VAE latents. This means fewer seams in generated images, more faithful text rendering, and better preservation of subject identity during editing. For product builders, this translates to more reliable outputs and less need for post-processing fixes.

  2. Spatial coherence is a differentiator. The spatially coupled decoder is the paper’s technical centerpiece, and the results suggest it is a meaningful advance. The ability to generate at 4K resolution with consistent texture and geometry across patch boundaries opens the door to high-fidelity applications—from detailed architectural renderings to scientific visualization—where previous encoder-VAE systems would have struggled with visible artifacts.

  3. Generation as reasoning. The finding that SenseNova-U1.5's generative capabilities enhance its reasoning (and vice versa) is subtle but important. It suggests that in a native unified model, "seeing" and "creating" are not opposite poles but reinforcing capabilities. A model that can generate an image of a concept is often better at reasoning about that concept, and a model that can reason about an image can generate more precise depictions. This blurs the boundary between perception and creation in a way that could spawn new types of interactive AI.

  4. Watch the structured instruction generalisation. The paper notes that despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions. This hints at a broader principle: multimodal understanding and reasoning can transfer to visual planning and creation. As the field moves toward models that can follow nuanced, multi-step visual directives—imagine an AI assistant that can "redesign this floor plan to improve flow while keeping the kitchen triangle"—this transfer capability will be a key differentiator.

SenseNova-U1.5 is a strong demonstration that the "native unified" approach is more than a philosophical stance. It is a practical architecture that scales, generalizes, and delivers state-of-the-art results across generation, editing, and reasoning. For anyone building or evaluating multimodal AI, it offers a compelling new reference point: a model that can look at an image, reason about it, and create new visual content—all within the same framework. The open-sourcing of their training code means the research community can audit, extend, and build upon these ideas, which is perhaps the most valuable contribution of all.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →