arXiv:2608.23392nvidia/nemotron-3.5-lightning-30b-a3bAugust 24, 2026

Towards a Densing Law for User Representation Learning at Billion-Scale CapacityExplained for Beginners

Bin Dou, Junru Zhang, Zhaoyi Yuan +6 more

Information RetrievalArtificial Intelligence

Abstract

User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance exhibit diminishing performance gains with larger-scale raw text user behavioral input, which can be mitigated by tokenization. (ii) Lack of quantitative analysis of how tokenization configurations should scale with data size. In this report, we propose User Behavioral Densing Law for characterizing the quantitative relationship between data scale and the minimum sufficient tokenization capacity. Firstly, we conduct a pilot study on raw & tokenized scaling comparison on billion-scale Alipay dataset, revealing the raw data scaling bottleneck and the sustained gains enabled by tokenization. To derive the scaling pattern governing the minimum sufficient tokenization configuration at different data scales, theoretical analysis and systematic experiments are employed to summarize the quantitative scaling pattern. We find an approximately linear relationship between the logarithms of minimum sufficient tokenization capacity and input data size measured by tokens, and the scaling slope varies systematically with the tokenization method and data source, reflecting differences in representation-space redundancy and intra-source uniqueness. Guided by the proposed law, we further develop ALGN, an adaptive variable-length tokenization method that improves capacity allocation. Extensive experiments across diverse data sources, tokenization methods, and downstream tasks demonstrate the generalizability and reliability of the User Behavioral Densing Law, providing practical guidance for tokenization configuration selection in large-scale user representation learning. Moreover, ALGN outperforms existing baselines.

Here is the structured explanation based on the provided paper.

1. The Problem

The core problem this paper addresses is the scaling bottleneck in industrial user representation learning. Modern platforms like Alipay rely on user embeddings to power everything from recommendations to risk control. Traditionally, the way to make these systems better has been to throw more data at them—more users, longer behavioral histories, and larger AI models.

However, the authors argue this approach is hitting a ceiling. In real-world behavioral data, volume grows much faster than useful information. User behavior is heavily shaped by daily routines and repeated habits; as you add more users or more days of data, the "signal" (task-relevant information) grows slowly, while the "noise" (redundant, repetitive events) explodes. The paper coins this the "raw behavioral scaling wall." Simply feeding more raw text into a model yields rapidly diminishing returns on downstream tasks like click-through rate prediction or user targeting.

Furthermore, there is a lack of quantitative guidance. Practitioners know that tokenization—compressing raw behavior into discrete tokens—helps, but there hasn't been a rule of thumb for how to configure the tokenization as data grows into the billions. You can't just keep increasing token counts haphazardly; you need to know the minimum amount of "capacity" required to preserve performance.

Why should anyone outside academia care? Because this isn't just an academic exercise in compression. For a product manager or engineer at a tech company, this means the "free lunch" of just buying more GPUs and collecting more logs is over. To keep improving personalization features, you need smarter ways to process the data you already have. This paper provides the blueprint for that.

2. How It Works (The Technical Mechanics)

The paper proposes shifting from "volume scaling" to "density scaling." Instead of feeding long, repetitive raw sequences into the encoder, the system first transforms them into compact discrete tokens using a method called Residual Quantized Variational Autoencoder (RQ-VAE).

The Mechanics in Plain English: Think of a user's behavioral history as a long playlist of songs. If you play that playlist for an AI model raw, the model wastes a lot of "brain cells" remembering that the user listened to the same song yesterday, and the day before, and the day before that. It’s redundant.

The tokenizer acts like a smart summarizer. It looks at that playlist, identifies the "essential" tracks (high-entropy, unique events), and compresses the rest into a short code. In the paper's specific implementation, they use RQ-VAE, which works in stages: it first codes the basic events, then codes the "residual" (the error) using additional codebooks. This results in a sequence of discrete tokens (like integers ID1, ID2, ID3) rather than continuous floating-point numbers.

The Analogy: The "Playlist" vs. The "Setlist" Imagine you have a user who listens to 100 songs a day, but 90 of them are just "commute to work" or "buy coffee." If you feed 30 days of this raw data (3,000 songs) to a model, the model spends capacity learning that "commute" is a thing.

Now, imagine a tokenizer that says, "I'll compress 3,000 songs into just 50 tokens." It keeps the 30 unique, interesting songs and summarizes the 2,970 repetitive ones into a few codes like [repeated routine] and [repeated purchase]. The AI model now has to process only 50 tokens instead of 3,000. It spends its capacity learning what makes this user unique rather than how often they buy coffee. This is the "densing" effect—increasing the information density per unit of input.

Defining Key Terms Inline:

  • Tokenization: The process of transforming long raw histories into compact sequences of discrete codes.
  • Residual Quantization (RQ-VAE): A technique that recursively quantizes residual information using multiple codebooks, allowing for multi-level discrete representations.
  • Information Density: The amount of task-relevant signal packed into each input unit. The paper argues this is the real bottleneck, not model size.

3. Key Results & Benchmarks

The experiments were conducted on a massive Alipay dataset (billions of users) and evaluated on three downstream tasks: Classification (advertising/marketing), Text-based Retrieval (user targeting), and U2U Retrieval (few-shot user targeting).

The Raw Data Wall: The paper visually demonstrates the "saturation" point. In Figure 4, when they increase the user population from 0.01B to 0.1B (while fixing other things), performance jumps initially but then flattens completely. Past 0.03B users, adding more data does practically nothing for accuracy. The same happens when extending the behavioral window from 60 to 120 days—you double the input, but gain very little in return.

The Tokenization Breakthrough: When the same models were fed tokenized data (using RQ-VAE), the results flipped. Figure 5 shows that tokenized representations maintain consistent performance gains even as model size scales up from 0.05B to 0.4B parameters. Crucially, the tokenized models outperformed the raw data models at every scale.

Concrete Impact Translation: The paper reports that tokenized data overcomes the "raw scaling wall." For example, in the classification task, tokenized models achieved higher AUC scores than raw models even with significantly smaller input budgets. The authors state that under matched token and compute budgets, tokenized representations "overcome the raw scaling wall" and yield "consistent performance gains and improved encoder utilization."

While exact percentage improvements vary by task, the consistent takeaway is that tokenization converts the curve from a flattening line (diminishing returns) back into an upward slope (sustained gains).

4. Why It Matters (Key Takeaways)

Here are the four most important takeaways from the paper:

  • The Shift to Density Scaling: The era of "the more data, the better" is over for raw behavioral logs. To scale further, engineers must focus on increasing the information density of the input, not just the volume. If you are collecting petabytes of logs, a significant portion is likely redundant.
  • Tokenization as a Performance Multiplier: Tokenization isn't just about saving memory; it's a way to recover performance lost to redundancy. The paper shows that properly tokenized data allows models to achieve better results than raw data, even when the raw data is vastly larger.
  • The "Densing Law" as a Configuration Guide: The authors propose a rough linear relationship between the logarithm of the minimum token count and the logarithm of the data size. In practical terms: as your user base grows 10x, you don't need 10x more tokens. You need a calculated, smaller increase. This gives product teams a quantitative target for how large their codebooks or token budgets should be.
  • The Advent of ALGN: The paper introduces ALGN (Adaptive Length Gated Network), a new tokenizer that dynamically allocates more "capacity" (codes) to erratic, complex user behavior and fewer codes to predictable routines. ALGN outperforms existing fixed-tokenization baselines, proving that adaptive, variable-length tokenization is the future for billion-scale systems.

What to Watch For Next: The industry will likely see a move away from fixed-length embedding layers toward these adaptive tokenization schemes. The key limitation to watch is the added complexity of the tokenizer itself—while it saves compute downstream, the process of generating those discrete tokens requires its own engineering overhead. Additionally, as with any compression technique, there is a risk of "over-compressing," where unique but rare user behaviors get smoothed over into generic codes.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →