arXiv:2608.27875nvidia/nemotron-3.5-lightning-30b-a3bAugust 28, 2026

HyQuant: Hybrid-Precision Quantization for LLM AttentionExplained for Beginners

Jiatong Ding, Bingxin Xing, Yu Zhang +9 more

Artificial Intelligence

Abstract

Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the attention module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose HyQuant, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .

The Problem: The Quantization Accuracy Cliff

Imagine you’re trying to stream a high-definition movie, but your internet bandwidth is limited. To fit the data through, you compress the video. However, if you compress too aggressively—say, reducing the color depth to just 4 bits—the picture breaks down. Faces become blurry, motion artifacts appear, and you lose the details that matter for understanding the story.

This is the core problem HyQuant addresses in Large Language Models (LLMs). Quantization is the process of reducing the precision of a model’s numbers (e.g., from 16-bit floating point down to 4-bit integers) to make the model smaller, faster, and cheaper to run. It has been a game-changer for deploying LLMs on consumer hardware.

However, the attention mechanism—the part of the LLM that decides which parts of the input text are most relevant for the current output—is notoriously fragile under this compression. In low-bit quantization, errors are not distributed evenly. A few "outlier" tokens, particularly those forming persistent vertical-line patterns in the attention heatmap, dominate the model's output. If these critical tokens are crudely quantized to 4 bits, the model’s reasoning ability crashes.

Existing methods have largely tried to smooth out these errors or treat all tokens as equally important. But as the paper notes, this is a "one-size-fits-all" approach that wastes precision on forgettable tokens while starving critical ones. The gap HyQuant fills is designing a system that recognizes which tokens are the "stars" of the attention show and treats them differently from the "background extras."

How It Works: The Hybrid-Precision Strategy

HyQuant operates on a simple but powerful insight: attention is non-uniform. A small fraction of tokens (typically less than 5%) are repeatedly attended to across long generations. These form vertical stripes in the attention matrix—hence the name "vertical-line tokens." They act as anchors for the model's reasoning.

The framework splits the workload into three groups:

  1. Vertical-Line Tokens (KVLK_{VL}): These are the persistent, high-attention positions. HyQuant keeps these in full precision (e.g., 16-bit).
  2. Local Window (KWinK_{Win}): A recent sliding window of tokens (e.g., the last 128 tokens). These are also kept in full precision because recent context is crucial for coherence.
  3. The Quantized Majority (KQK_Q): Everything else is quantized to low bit-widths (4-bit for both keys and values).

The Technical Mechanics: Two Operational Modes

  • Prefill Stage (Compute-Heavy): During the initial processing of a long prompt, the model does massive matrix multiplications. HyQuant introduces a "fused operator." Imagine a conveyor belt where most items are moving at a fast, low-resolution setting (INT4), but a few critical items are moved separately at high resolution (FP16). The operator scans the quantized majority first, then overlays the high-precision vertical lines and local window into the result. Because the critical tokens are few, the overhead of keeping them precise is tiny, but the accuracy gain is massive.
  • Decode Stage (Memory-Heavy): Once the model starts generating tokens one by one, the bottleneck shifts to memory bandwidth—the speed at which the model can read past key-value states from memory. HyQuant compresses the KV cache (the model's memory of past tokens) into 4-bit. However, it keeps the vertical-line tokens and the local window in full precision. Crucially, it fuses the dequantization step into the attention computation itself. Instead of reading a 4-bit cache, decompressing it to 16-bit in a separate step, and then computing attention, HyQuant decompresses on the fly during the computation. This saves a significant amount of memory traffic.

Analogy for the Engineer: Think of it like a database query. Normally, you might compress all data entries to save disk space (uniform quantization). HyQuant is like a smart query optimizer that identifies the "primary keys"—the records accessed most frequently by every other query. It keeps those primary keys in full, uncompressed detail, while compressing the millions of rarely-accessed log entries to a compact format. The result is a smaller database that doesn't slow down your standard lookups.

Key Results & Benchmarks: Nearly Lossless, Much Faster

The results are striking because they decouple accuracy from efficiency. Often, pushing for speed means sacrificing some intelligence. HyQuant bucks this trend.

The Numbers:

  • Speed: On the decode stage, HyQuant achieves a 1.32× to 3.58× speedup over the full-precision FlashAttention-2 baseline as context length grows to 32,768 tokens. This is a huge win for long-context tasks.
  • Accuracy: On LongBench (a suite of long-context reasoning tasks), HyQuant maintains accuracy comparable to the full-precision baseline. In some math reasoning benchmarks (GSM8K, MATH500), it actually slightly edges out the full-precision model, which the authors attribute to evaluation variance rather than genuine improvement.
  • Baseline Comparison: Compared to strict 4-bit baselines like KIVI or KV-Tuner, HyQuant is consistently better. For instance, on the LongBench average, KIVI might score around 82, while HyQuant pushes close to 88 or higher, depending on the model.

Translating Benchmarks: Consider the GSM8K benchmark, which tests mathematical reasoning. The full-precision model (FA2) scores about 95.88. The strict 4-bit KIVI baseline drops to about 92.48—a decline of roughly 3.4 points. HyQuant, with its hybrid approach, scores 96.52. In plain language, this means HyQuant doesn't just prevent the crash; it arguably preserves more of the model's original mathematical capability than the standard 4-bit approach, all while using the same or less compute.

Why It Matters: Key Takeaways

  1. The "Vertical Line" is the Anchor: The paper provides rigorous evidence that attention is concentrated. Roughly 5% of tokens cover 50-85% of the attention mass. Ignoring this structure is the primary reason low-bit quantization fails. HyQuant’s success proves that "sensitivity-aware" compression is the right direction.
  2. The Fused Kernel is the Secret Sauce: The performance gains don't come from the quantization scheme alone, but from the engineering of the operator. By fusing dequantization with attention computation, HyQuant avoids the "memory wall"—the latency penalty of shuttling data between memory and compute units.
  3. Practical Overhead is Minimal: The authors measure the extra cost of identifying which tokens are vertical lines at only 3% to 5% of total runtime. This is a negligible tax for the accuracy and speed benefits gained.
  4. The Memory Trade-off: There is a memory cost. Retaining 5% of tokens in full precision and a local window increases the KV cache size by about 15-24% compared to strict 4-bit quantization. However, the end-to-end speedup (1.04× to 1.17×) and the massive reduction in quantization error make this trade-off worthwhile for long-context scenarios.

What to Watch For:

  • Short Context: The benefits diminish in very short contexts (under 1K tokens), though the model remains competitive.
  • The 5% Ratio: The paper uses a top-5% threshold. If the task changes or the model architecture shifts, the optimal ratio might need tuning.
  • Hardware Specificity: The speedups were measured on NVIDIA H100 GPUs. On other hardware, the exact balance of compute vs. memory bandwidth benefits might shift.

Summary: HyQuant is a pragmatic solution to a persistent problem in LLM deployment. It moves away from the fantasy of "uniform quantization"—where every number is treated equal—and embraces the reality that attention is lopsided. By reserving the "highway lanes" for the critical tokens and compressing the side streets, it achieves a rare trifecta: significant speed gains, massive memory savings, and negligible accuracy loss.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →