arXiv:2609.04010nvidia/nemotron-3.5-lightning-30b-a3bSeptember 3, 2026

Unlocking Lossless Speedups in LLMs via Discrete DiffusionExplained for Beginners

Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham +14 more

Machine Learning

Abstract

Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weights are learned through a simple Diffusion Distillation phase that adds negligible overhead to existing LLM training pipelines. We also introduce Ψ-Spec, a family of samplers that enables lossless acceleration and inference-time scaling at a fixed context length. Unlike speculative decoding, our method requires no separate draft model. Unlike diffusion LLMs (d-LLMs), it accelerates generation without sacrificing the quality of the underlying AR model. The resulting models, called Uno, can be trained from scratch or built by augmenting existing open-weight AR LLMs. Uno achieves higher throughput than leading speculative-decoding methods at every evaluated batch size and delivers up to 3times speedups over the base AR model, including at the largest batch size supported by the device. Notably, our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning. We release code and checkpoints at: https://s-sahoo.github.io/uno/

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

The Problem: Autoregressive LLMs Are Too Slow

Here’s the bottleneck that every AI product manager has felt: Large Language Models (LLMs) are built on next-token prediction (NTP). You give the model one token, it predicts the next one, you feed that back in, and repeat. This autoregressive (AR) structure is why generating a thousand tokens takes a thousand sequential steps.

That’s painfully slow for two reasons:

  1. Sequential dependency – You can’t predict token 3 until tokens 1 and 2 are done. This kills latency for long responses and makes reinforcement-learning rollouts (where the model generates many trial sequences) incredibly expensive.

  2. Memory bandwidth limits – Modern GPUs are great at parallel math operations, but LLM inference is often memory-bound. Moving model weights and key-value states in and out of GPU memory is the real speed limit, not the computation itself.

Language actually contains lots of predictable patterns—formulaic sequences, common phrases, code scaffolding—but standard LLMs can’t exploit this redundancy because they’re forced to generate one token at a time.

Existing approaches only partially fix this. Speculative decoding uses a separate “draft model” to guess multiple tokens, then verifies them with the big model. It works, but requires maintaining an extra model and its gains depend on having a well-aligned drafter. Discrete diffusion models support parallel generation natively, but the leading diffusion-based LLMs (d-LLMs) still face a quality-versus-speed tradeoff, and their speedups vanish at large batch sizes—the very regime where real-world serving happens.

What gap does this paper fill? The authors ask: Can we get the parallelism of diffusion models while preserving the quality of autoregressive models—and do it without training a separate draft model or sacrificing quality at scale?

How It Works: The “Uno” Architecture

The paper introduces diffusion-augmented LLMs, a framework that adds a second pathway through each transformer layer: diffusion weights trained to generate tokens in parallel, while the original AR weights preserve response quality.

Two Decoupled Weight Sets

The key insight is parameter decoupling:

  • AR weights (θAR) – Trained via the standard NTP objective. These determine response quality. They’re frozen during the diffusion training phase.

  • Diffusion weights (θΔ) – Lightweight LoRA adapters added to each layer. Trained exclusively to generate token blocks in parallel. They add minimal memory overhead (rank-128 LoRA adapters with α=256).

This separation means you can train the AR weights once using the established pipeline (pre-training, SFT, RL post-training), then add diffusion weights on top with negligible overhead. The two pathways are coupled through gated LoRA: the draft pathway uses both θAR and θΔ, while the verification pathway uses only θAR. This coupling ensures the diffusion drafts stay aligned with the AR verifier without requiring a separate model.

The Training Pipeline

The training has two stages:

  1. NTP Training – Standard autoregressive training on massive text corpora. This produces a high-quality AR model.

  2. Diffusion Distillation – Freeze the AR weights, then train the diffusion LoRA adapters to reverse a uniformly-corrupted sequence in a single step. The loss has two terms:

    • LDCD (Distillation Loss) – Matches the diffusion output to the AR distribution over token blocks.

    • LTV (Total Variation Loss) – Encourages consecutive diffusion tokens to be accepted by the AR verifier, increasing draft acceptance length.

    Training uses a blockwise approach: the sequence is partitioned into blocks, and each block is denoised conditioned on preceding context. The authors train on 7B tokens with block sizes of 2, 4, and 8, using a global batch size of 128.

Why this matters: The diffusion distillation phase adds essentially zero overhead to existing LLM training pipelines. You could imagine it as an optional add-on after standard fine-tuning.

The Ψ-Spec Sampler

At inference time, the Ψ-Spec sampler uses the diffusion pathway to propose a block of tokens in parallel, then applies speculative-decoding-style rejection sampling against the AR model to accept the longest valid prefix. The sampler has two configurations:

  • Linear Sampler – Samples one candidate per token position. Best for high batch sizes where the system is compute-bound.

  • Tree Sampler – Samples multiple candidates (top K at each position) and verifies them in a prefix tree. Best for low batch sizes where there’s spare compute capacity.

A crucial property: the sampler preserves the AR model’s output distribution exactly. This means the speedup is lossless—you get faster generation without degraded answer quality.

Key Results & Benchmarks

The paper evaluates Uno (their diffusion-augmented models) against multiple baselines across agentic tasks, math, coding, and long-context reasoning.

Throughput Gains

ModelSystem Throughput (tokens/sec) at Largest Batch
Uno (8B)5,255
Base AR model3,522
DiffusionGemma (26B)1,136
Nemotron-Labs-Diffusion (14B)1,197
Mercury 2 (proprietary)1,154

At the largest batch size supported by the base AR model (batch 64), Uno is 1.5× faster. At batch size 1, Uno is ~2.2× faster. These aren't marginal improvements—they're order-of-magnitude shifts in what's practically possible.

Quality Preservation

The paper emphasizes that Uno maintains the AR model's output distribution. From the results:

  • Uno vs. base AR: "Uno qualitatively matches the base AR model while achieving higher throughput at every batch size."
  • After RL post-training: Diffusion adapters trained on the SFT checkpoint "retain their speedups after RL post-training, with only a nominal 6% decrease in TPFs."

This lossless property is what separates Uno from d-LLMs like DiffusionGemma, which the paper notes "reports a system throughput of 1,154 tokens/s for a 1K-token input at batch size 10. Uno achieves a maximum system throughput ∼4.6× higher, despite Mercury 2 running on substantially faster Blackwell GPUs."

Comparative Benchmarks

The 8B Uno model outperforms:

  • DiffusionGemma (26B) – on all evaluated benchmarks
  • Nemotron-Labs-Diffusion (14B) – on all evaluated benchmarks
  • Mercury 2 (proprietary) – on agentic tool use, coding, and long-context reasoning

Notably, this is an 8B model beating a 26B model and a proprietary model that runs on faster hardware. The speed-quality tradeoff is genuinely broken.

RL Training Speedup

When used for reinforcement learning rollouts, the diffusion adapters trained on the SFT checkpoint provide "up to a 40% end-to-end training speedup, primarily for the mathematics and code experts." Even after consolidating multiple experts via ISO-Merger, the speedups persist with only a 6% TPF decrease.

Why It Matters: Key Takeaways

  1. Lossless speedup at scale – Uno achieves up to 3× speedup over the base AR model across all batch sizes, including the largest batch that fits on device. Previous methods' speedups typically diminished at high concurrency; Uno's persist.

  2. No separate draft model needed – Unlike speculative decoding (EAGLE-3, DFlash), Uno uses a single architecture with dual pathways. This means:

    • No extra model to maintain
    • Single KV cache (lower peak memory)
    • Fewer additional parameters (0.35B vs. 1.05B for DFlash)
  3. Drop-in augmentation for existing models – Uno can be created by augmenting any open-weight AR LLM with diffusion weights. The paper demonstrates this with Qwen3-8B, showing that even when the AR weights are fine-tuned on a different distribution (OpenThoughts), the diffusion weights still enable lossless acceleration.

  4. Practical impact across workloads – The speedups benefit both inference serving and RL post-training. For agentic workloads (where a single request triggers parallel agents, tool calls, and retries), the system throughput gains are substantial. For RL training, 40% faster rollouts means cheaper, faster model development.

  5. What to watch – The method works best when the diffusion and AR weights are reasonably aligned. The gated LoRA coupling helps, but extreme distribution shifts between the diffusion training data and the AR training data could eventually degrade acceptance rates. The paper's open-weight experiments (Qwen3-8B + OpenThoughts) show this is manageable in practice, but it's worth monitoring as the approach scales to larger models and more diverse domains.

Bottom line: This paper delivers on a long-sought goal in LLM inference—parallel token generation without quality loss. By decoupling AR and diffusion weights within a single architecture and using a clever distillation + rejection-sampling pipeline, Uno achieves real speedups that compound under real serving conditions (high batch sizes). The fact that an 8B model beats 26B and proprietary alternatives across diverse benchmarks suggests this isn't just a laboratory curiosity but a genuinely better way to build and run LLMs.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →