arXiv:2608.25936nvidia/nemotron-3.5-lightning-30b-a3bAugust 26, 2026

One Symptom, Three Levers: A Critical Review of On-Policy Self-DistillationExplained for Beginners

Justin Robert, Raheel Qader

Machine LearningArtificial IntelligenceComputation and Language

Abstract

On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were promising, with accuracy comparable to reinforcement learning at a fraction of the generated tokens. But the same asymmetry that produces the signal also biases it. One failure mode now dominates the field: collapse, the progressive narrowing of the set of reasoning paths the model can produce. Collapse is not specific to OPSD, though privileged information aggravates it. This review treats collapse as a symptom governed by three levers: (i) where the signal is applied, that is, how tokens are weighted; (ii) what the teacher is shown, that is, the nature of the privileged information; and (iii) when the signal changes, that is, the teacher's dynamics and the decay of guidance. We restrict our scope to mathematical reasoning, where the method originated and where its failure modes are best documented. We report no new experiments. The contribution is structural: a shared vocabulary for phenomena named differently across papers, and a clear line between what is settled and what is still disputed.

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

In the race to make large language models reason, a new contender has emerged that promises the best of both worlds: the dense supervision of imitation learning without the cost of a separate, larger teacher model. Called On-Policy Self-Distillation (OPSD), the method has generated significant excitement since its proposal in early 2026. But as a recent critical review argues, the very asymmetry that makes OPSD possible also makes it fragile. The paper treats "collapse"—the progressive narrowing of the model's reasoning abilities—not as a bug, but as a symptom governed by three controllable levers.

The Problem: Efficiency vs. Diversity

To understand why OPSD matters, we first have to look at the landscape it entered. For a while, Reinforcement Learning with Verifiable Rewards (RLVR) was the dominant route to reasoning ability. In RLVR, a model generates several "rollouts" (attempts at solving a problem), and a reward is scored only at the very end based on whether the final answer is correct. While this avoids the need for a second model, it has structural limits. The reward is sparse—a single signal must account for hundreds of tokens, making it difficult to know which specific reasoning steps led to success. It is also expensive, requiring many long rollouts per problem, and tends to concentrate probability mass on the reasoning the model already knows, rather than discovering new paths.

Enter On-Policy Distillation (OPD), which tried to restore a dense token-by-token signal without abandoning on-policy sampling. The price was a second, larger model to act as the teacher. On-Policy Self-Distillation (OPSD) sought to remove that cost. Instead of a larger model, the teacher is the model itself, but conditioned on "privileged information"—a reference solution, a plan, or environment feedback—that the student will not have at test time.

Early results were promising. OPSD matched or exceeded the performance of RLVR on mathematical reasoning while generating far fewer tokens. It offered autonomy: no larger model was required. But the paper makes a sobering observation. The same asymmetry that produces the dense signal also biases it. The dominant failure mode now is collapse: the progressive narrowing of the set of reasoning paths the model can produce. The model might get more problems right on average, but it loses the ability to explore rare but correct solutions. It becomes brittle, performing well on training distribution problems but poorly on new ones.

How It Works: The Mechanics of Self-Distillation

The technical core of OPSD is a loop that mirrors Generalized Knowledge Distillation (GKD), but with a crucial twist. In standard distillation, a large "teacher" model guides a smaller "student." In OPSD, there is no second model. The student generates a rollout (an answer), and then the teacher—which is the same model, but given extra information—scores it token by token.

Here is the step-by-step mechanics:

  1. Generation: The student model is prompted with a problem xx. It samples an answer autoregressively, token by token.
  2. Scoring: The "teacher" (the same model architecture) scores the student's rollout. Critically, the teacher is given the problem xx plus the privileged information y∗y^* (e.g., the reference solution). For each token the student generates, the teacher predicts the next token probability distribution, but it conditions this prediction on both the problem and the reference solution.
  3. Divergence: The method computes the gap between the teacher's distribution and the student's distribution at each token position. The founding paper uses the forward KL divergence. Think of this as the teacher telling the student, "Cover all the possibilities I think are valid." This is distinct from the reverse KL, which would push the student to pick the single most likely token the teacher chose.
  4. Learning: The gradients flow only into the student's weights. The teacher remains fixed at its initial policy.

The efficiency gain is real. Where RLVR might sample eight rollouts of up to 16,000 tokens per problem, OPSD makes do with a single generation capped at 1,024 tokens. At comparable performance on math benchmarks, it consumes far fewer generated tokens per problem. However, the paper notes this does not translate to a compute gain per step; an OPSD step requires two forward passes and one backward pass, roughly twice the cost of a GRPO step. The advantage is faster convergence in the number of steps, not a lower unit cost.

Key Results and Benchmarks

The results are compelling but fragile, as the review emphasizes. On mathematical benchmarks like AIME25, OPSD shows strong gains. The founding paper reports a rise from 36.7% to 43.9% accuracy on AIME25 by step 50, using the forward KL divergence. Best reported scores for a 1.7B model reached 57.2% on AIME24 and 43.9% on AIME25. However, the review warns that these numbers are "best-over-checkpoints" figures; the end-of-training value on AIME25 was 41.1%.

The review is highly critical of how these numbers are typically reported. It identifies three sources of illusion that plague small-model benchmarks:

  • Variance: AIME has only 30 questions; a single question flipping shifts the score by more than three points.
  • Contamination: AIME 2024 problems are partly present in pre-training data, meaning some models complete them from memory.
  • Model-family specificity: On Qwen models, even a random training signal can raise scores due to pre-training artifacts, an effect absent in other families like Llama or OLMo.

Furthermore, the review highlights a disturbing trade-off. While OPSD improves average accuracy, it reduces diversity. The paper cites work showing that pass@1 rises while pass@k flattens or declines. A model might succeed more often, yet lose the ability to explore rare but correct solutions. On some models, token entropy actually increases (suggesting more randomness) while functional diversity decreases (fewer distinct reasoning paths). This means relying on mean score alone can mask a catastrophic loss of capability.

Why It Matters: The Three Levers

The heart of the paper is its structural analysis of why collapse happens and how to control it. The authors argue that collapse is not a single thing, but a symptom governed by three levers. Understanding these levers is essential for anyone looking to use or build upon OPSD.

Lever A: Signal Geometry (Where and how the signal is applied)

The first lever concerns the "shape" of the distillation signal. The method must choose a divergence measure (forward KL, reverse KL, or JSD) and decide on the "density" of the signal (which tokens are supervised and how heavily).

  • The Direction: The choice of divergence is critical. The reverse KL pushes the student to imitate the teacher's most confident prediction (mode-seeking), which can dramatically reduce diversity. The founding paper chose the forward KL, which is "mass-covering" and preserves coverage of the teacher's behaviors. However, even forward KL can suffer if not weighted properly.
  • The Density: The intuition that "denser is better" is directly contradicted by recent work. Distilling the full chain of thought helps on short tasks (like tool use) but degrades mathematics and science, whose long traces surface artefacts. The review points to "selective density" as the promising path: weighting tokens by importance. For instance, Entropy-Aware OPD applies reverse KL generally but adds forward KL only on the teacher's high-entropy tokens (where the teacher is uncertain). This preserves diversity where the signal is ambiguous. Another approach, DPH-RL, partitions problems: for "mastered" problems, it uses a mass-covering divergence to anchor the model; for "unmastered" problems, it removes the divergence entirely, allowing free exploration.

Lever B: The Nature of Privileged Information (What the teacher is shown)

This is perhaps the most practically significant lever. The paper argues that not all privileged information is created equal, and the kind of information given to the teacher dictates whether the student learns useful skills or dangerous shortcuts.

The review offers a taxonomy ordered by risk, based on how much the information presupposes the solution:

  • The Final Answer (Oracle): The riskiest. A teacher that already knows the answer expresses almost no uncertainty, pushing the student to skip reasoning steps and take shortcuts. It is "inert" or even harmful in self-distillation regimes.
  • The Worked Solution (OPSD default): The reference derivation plus the answer. It presupposes the answer as fully as the oracle.
  • The Full Chain of Thought (CoT): Exhibits the reasoning but may not state the final answer. Showing only the first half of the CoT is often preferable.
  • The Plan/Skeleton: Gives the structure without the values. This is more reachable for the student.
  • The Rubric: Lists criteria for a good answer without imposing a specific path, preserving diversity.
  • Error Feedback: Partial and conditional information (e.g., error messages). The least presupposing.

Two recent studies are particularly illuminating. One contrasted an aligned critique (which corrects only the student's incorrect steps) against the reference solution. The aligned critique won by a significant margin, modifying learning signals only where the reasoning fails, rather than at every token. Another study found that step-wise hints without execution dominated the final answer as privileged information. The takeaway is clear: privileged information helps only if the student can reconstruct it at test time. If the teacher's information is too prescriptive, the student memorizes shortcuts it cannot replicate independently, degrading out-of-domain performance.

Lever C: Loop Stability (When the signal changes)

The final lever deals with temporal dynamics. In OPSD, the teacher is frozen at the initial policy. As the student learns and its weights change, the gap between the student and the teacher can widen, or the teacher can become obsolete.

The review draws parallels to self-supervised learning in vision, where identical loops were solved five years ago with three rules:

  1. Stop-Gradient: Never backpropagate the gradient into the teacher. Its output is treated as a fixed target. This is the "decisive ingredient" to prevent collapse.
  2. EMA Teacher: Let the teacher evolve, but slowly, as an exponential moving average of the student's past weights. The teacher "lags behind" the student, providing stable yet improving targets.
  3. Structural Asymmetry: The student must be structurally different from the teacher (e.g., the teacher sees privileged information the student does not). This imbalance prevents the two models from merging into a single, degenerate mode.

A second, independent lever is the decay of guidance. Rather than showing the teacher the full reference trace from the start, the share of the trace revealed can be reduced over training. If the student learns quickly, less help is shown; if it stalls, more is revealed. This "adaptive exposure" has shown gains ranging from +0.95 to +2.33 in average accuracy.

Synthesis and Conclusion

The review concludes that the initial bet of OPSD—that density alone was the decisive variable—has shifted. Density is no longer the variable that matters most; what matters is where the signal is applied, what it encodes, and when it is allowed to change.

The paper delivers three established results:

  1. Privileged information helps only if the student can reconstruct it. The reference solution, the most natural choice, is among the least transferable.
  2. Collapse is not visible in mean score or entropy. A model can have higher entropy than an RL-trained one while producing fewer distinct reasoning paths. Only pass@k reveals the narrowing of capabilities.
  3. Everything turns on how the asymmetry is controlled.

Should OPSD be used in production today? The answer is a qualified "no." It remains a research technique with well-documented failure modes: poorly chosen privileged information degrades performance, and in continual learning, the dense version can forget more than standard RL. However, the method is genuinely appealing for organizations seeking to control the production of their own small models, provided it is used with the safeguards outlined in the paper—essentially functioning as a frugal post-SFT fine-tuning method rather than a turnkey training solution.

The contribution of this paper is structural: it provides a shared vocabulary for phenomena named differently across papers and draws a clear line between what is settled and what is still disputed. Six months of work have qualified the principle that a model can guide itself, provided the teaching side holds information the answering side does not—but everything turns on how that asymmetry is controlled.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →