arXiv:2608.17336nvidia/nemotron-3.5-lightning-30b-a3bAugust 18, 2026

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference AccelerationExplained for Beginners

Hanzhi Zhang, Qiao Zhang, Qinglei Cao +4 more

Artificial Intelligence

Abstract

Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial precision routing over hardware-aligned score tiles outside fused dense attention. We introduce TileMix, a tile-centric precision-routing kernel that makes numerical precision an executable spatial decision over score-tile groups within fused dense attention. TileMix partitions the attention matrix into hardware-aligned score tiles, packs routing decisions into compact bitmasks, and dispatches each tile group through FP16 or INT8 score computation while both paths update a shared online-softmax state. Scalable precision grouping lets each routing bit govern multiple adjacent key tiles, preserving hardware-aligned compute tiles and compact metadata at long contexts. By routing all legal tile groups, TileMix preserves dense token connectivity, requires no training, and supports grouped-query attention, variable-length batches, and INT8 key/value caches. Across LongEval, LV-Eval, and A100 prefill benchmarks on LLaMA, Qwen, and Vicuna, TileMix recovers long-context quality lost under uniform INT8 and improves prefill throughput over FP16, yielding a controllable accuracy-efficiency frontier across model families. The implementation is available at https://github.com/HanzhiZhang-Ulrica/TileMix.

TileMix: Mixed-Precision Attention for Faster LLM Prefill

1. The Problem

When large language models (LLMs) try to read a long document, they face a fundamental physics problem: attention has to compare every query token against every key token. If a document has 32,000 words, the model computes roughly 32,000 × 32,000 = over 1 billion score interactions. This quadratic explosion makes long-context prefill—the initial reading phase—prohibitively slow and memory-hungry.

Researchers have tried three main strategies to fix this:

  • Low-precision quantization (e.g., INT8): Represents numbers with fewer bits to speed up computation and shrink memory. However, most prior work applies a single precision mode across the entire attention matrix. For long contexts, this is too coarse: some score tiles are critical (e.g., retrieving a specific fact), while others can tolerate lower precision.
  • Sparsity-based methods (e.g., BigBird, Longformer): Prune away "unimportant" token interactions to reduce computation. While efficient, they deliberately destroy the full attention graph, meaning the model can no longer see complete token connectivity.
  • IO-aware fused attention (e.g., FlashAttention): Smartly streams data through on-chip memory to reduce slow external memory traffic, but still uses one uniform arithmetic path per kernel invocation.

The gap: No existing method lets you spatially route different precision levels across the attention matrix while preserving the complete, dense token connectivity that long-context tasks rely on. TileMix was built to fill exactly that gap.

2. How It Works (The Technical Mechanics)

TileMix introduces a "tile-centric precision-routing kernel." Here is the core idea broken down into three mechanical steps:

The Tile Architecture

The attention matrix (which maps query tokens to key tokens) is divided into small, hardware-aligned rectangles called score tiles. On modern NVIDIA A100 GPUs, these tiles match the size of Tensor Core operations (typically 128×128 or 64×64 elements). This alignment ensures the GPU can feed data to its arithmetic units efficiently without stalling.

Packed Bitmasks: The "Spatial Decision Map"

Instead of choosing one precision mode for the whole matrix, TileMix uses a packed bitmask for each row of query tokens. Think of this as a binary barcode laid across the key tiles.

  • How it works: Each bit in the 64-bit mask corresponds to a group of key tiles. A 0 means "compute this group in FP16"; a 1 means "compute it in INT8."
  • The "Grouping" Trick: Extending a 64-bit mask across hundreds of thousands of key tokens would be unwieldy. TileMix solves this by letting one bit govern multiple adjacent key tiles (e.g., a bit might control a group of 4 or 8 tiles). This preserves the hardware-aligned compute tiles and keeps the metadata tiny—just a few bytes per query row—even at very long contexts.

The Dual-Path Engine

This is the ingenious part. During the inner loop of attention computation, TileMix reads the bitmask for the current key-tile group and dispatches the computation accordingly:

  1. FP16 Path: Standard floating-point matrix multiplication.
  2. INT8 Path: The queries and keys are quantized to INT8, processed through Tensor Core matrix multiplication (which is significantly faster), and then rescaled back to FP16 using per-block scale factors.

Critically, both paths feed into the exact same shared "online-softmax" state. This state tracks the running maximum and normalizer for the row-wise softmax. By updating a common state, TileMix avoids the numerical instability that typically mixes precisions, and it preserves the dense connectivity of full attention.

Analogy for Engineers

Imagine you are processing a massive spreadsheet with 1 million rows.

  • Sparsity is like deleting rows you guess are unimportant. You save time, but you might delete the one row containing the answer you need.
  • Uniform Quantization is like forcing every cell in the entire spreadsheet to use low-precision ink. It’s fast, but some critical calculations lose accuracy.
  • TileMix is like equipping each row with a tiny switch. If the row is a "critical" one (perhaps it contains a date or a name), the switch sends it through the high-precision FP16 path. If the row is mundane data, it takes the fast INT8 path. All rows are still processed; none are deleted. The clerk (the GPU) just uses different tools depending on the row's needs.

3. Key Results & Benchmarks

The paper evaluates TileMix across three major benchmarks and several model families (LLaMA, Qwen, Vicuna). The results demonstrate a compelling "accuracy-efficiency frontier."

Long-Context Quality Recovery

The uniform INT8 baseline (labeled "One" in the paper) suffers significant accuracy drops, especially as context length grows. For example, on the LV-Eval benchmark with LLaMA 3.2 3B at 64k tokens, uniform INT8 drops to ~5.42% accuracy, while FP16 stays near 7.75%.

TileMix recovers this lost quality. By intelligently routing specific tile groups to FP16—particularly those critical for the task—it brings accuracy back up to near-FP16 levels even while using significant INT8 compute. For instance, at 75% INT8 coverage, TileMix configurations sit consistently between the uniform INT8 baseline and the FP16 reference, often outperforming the "One" configuration by a meaningful margin.

Prefill Throughput Gains

On the A100 prefill benchmarks (Table 2 in the paper), TileMix demonstrably speeds things up over pure FP16 FlashAttention.

  • At 4k tokens: FlashAttention achieves ~14.33 K tokens/s. One (uniform INT8) reaches ~29.80 K tokens/s. TileMix (mixed precision) pushes past even One, reaching ~31.80 K tokens/s with the SpTrans75 configuration (75% INT8 routing).
  • Scaling up: At 8k tokens, the throughput gains persist, though the absolute numbers shift. The ordering consistently shows TileMix > One > FlashAttention in terms of throughput, while quality trends show TileMix recovering the quality lost by One.

Numerical Behavior

Table 3 reports the mean absolute deviation from a fixed FP16 reference. The results show that deviation increases with INT8 coverage, as expected. However, the paper highlights that coverage (e.g., 5%, 10%, 25%) acts as a practical numerical-control knob. At 8k tokens, the jump in deviation is notable between 5% and 10% coverage, but stabilizing after that, offering engineers a way to balance speed vs. precision.

4. Why It Matters (Key Takeaways)

  • Controllable Accuracy-Efficiency Trade-off: TileMix gives practitioners a dial to turn. You can ask for "75% INT8 speedup" and get a model that maintains significantly higher quality than the "One" uniform INT8 approach. This makes deployment on resource-constrained hardware (like edge devices or cost-sensitive cloud inference) much more feasible.
  • No Training Required: The routing policy is "static, data-free structured templates." This means you can plug TileMix into an existing LLM off-the-shelf without needing to fine-tune or re-train the model. It is a drop-in kernel optimization.
  • Hardware-Friendly Design: By respecting hardware-aligned tiles and using compact bitmasks, TileMix avoids the memory overhead that kills performance in other sparse or mixed-precision approaches. It specifically supports INT8 key/value caches, which is vital for the "decode" phase of inference where tokens are generated one by one.
  • Broad Compatibility: The paper notes support for Grouped-Query Attention (GQA) and variable-length batches. This means it works with the common architectural patterns found in modern models (like Qwen and LLaMA) without special casing.

What to watch for next: The paper implicitly opens the door to dynamic routing, where the bitmask itself could be learned or generated based on the content of the prompt (e.g., attending more precisely to the start of a document). Additionally, the numerical analysis suggests that finding the optimal "mix ratio" of FP16 vs INT8 tiles per layer or per task is a rich area for further research.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →