arXiv:2608.24696nvidia/nemotron-3.5-lightning-30b-a3bAugust 25, 2026

On-policy Distillation with Verifiable RewardExplained for Beginners

Wenze Lin, Jiale Zhao, Xitai Jiang +5 more

Machine LearningArtificial Intelligence

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.

The Problem: Two Paradigms That Should Cooperate, But Don't

To understand why this paper matters, we need to look at the two dominant approaches for improving large language models after they’ve been trained—the so-called "post-training" phase.

On one side, we have Reinforcement Learning with Verifiable Rewards (RLVR). This approach uses automated "verifiers"—essentially, computer programs that can objectively check if a model’s answer is correct, such as a math solver that checks a final answer or a compiler that checks if code runs. The model is rewarded for generating correct answers and penalized for wrong ones. The appeal is clear: the model optimizes directly for the task we care about. However, RLVR suffers from a critical flaw: sparse feedback. The model only learns whether the final answer was right or wrong. It gets no signal about whether its intermediate reasoning steps were good or bad. This makes it incredibly difficult for the model to know how to improve its thinking process.

On the other side, we have On-policy Distillation (OPD). This approach treats a stronger "teacher" model as a source of wisdom. The student model is trained to mimic the teacher’s output at every single token, effectively absorbing the teacher’s knowledge step-by-step. This provides dense feedback at every step of the reasoning process. But there’s a catch: OPD is purely distributional. It tells the student, "be like the teacher," but it doesn't care if the teacher’s answer was actually correct. If the teacher makes a mistake, OPD will try to reproduce that mistake. This limits the student’s performance to the teacher’s level; the student can’t surpass the teacher because it’s only learning to imitate, not to reason.

The paper’s core insight is that these two methods are complementary. RLVR provides the "what" (correctness), but lacks the "how" (step-by-step guidance). OPD provides the "how," but blindly follows the teacher’s "what," correct or not. The natural thing to do is combine them. However, previous attempts to combine them have been messy, often relying on weighted sums or heuristic rules that require tuning new hyperparameters—extra knobs to turn that add complexity and trade-offs.

How It Works: The ReLU Gate

The paper’s brilliance lies in its simplicity. Instead of treating OPD and RLVR as separate objectives to be balanced, it reformulates OPD to inherently respect task correctness.

Here is the mechanical core of the idea:

  1. The Problem with Standard OPD: In standard on-policy distillation, the reward for a token is essentially the log-ratio of the teacher's confidence to the student's confidence (log⁡(πT/πθ)\log(\pi_T / \pi_\theta)). The issue is that this sign depends on who is more confident, not whether the answer is right. If the student is overconfident on a wrong answer, the reward might be positive, accidentally reinforcing a mistake. If the teacher is more confident on a correct answer, the reward might be negative, punishing a correct behavior.
  2. The ReLU Gating Mechanism: OPDVR applies a Rectified Linear Unit (ReLU) gate to this reward. ReLU simply outputs the input if it's positive, and zero if it's negative. In this context, it acts as a sign corrector.
    • For correct trajectories: The reward is forced to be non-negative. If the student is already more confident than the teacher on a correct step, the gate zeros out the update (preventing the model from being "corrected" away from a working solution).
    • For incorrect trajectories: The reward is forced to be non-positive. If the teacher is more confident than the student on a wrong step, the gate ensures a negative reward, suppressing that erroneous path.

The paper uses a compelling analogy: Think of the teacher as a knowledgeable mentor and the verifier as a strict boss. Standard OPD listens only to the mentor, regardless of whether the boss is happy. OPDVR listens to the mentor for how to think, but constantly checks with the boss to make sure the reasoning is actually on the right track. If the mentor suggests a path the boss rejects, OPDVR ignores the mentor for that step. If the mentor suggests a great path the boss approves of, OPDVR reinforces it even more strongly.

Furthermore, because this modification transforms the OPD reward into a valid RLVR signal, it becomes compatible with standard policy gradient algorithms. The paper demonstrates this by integrating OPDVR with GRPO (Group Relative Policy Optimization), a popular RL algorithm, creating what they call Group Relative Policy Distillation (GRPD).

Key Results: Beating the Teacher and the Baseline

The experiments are the proof point. The paper evaluates the method on six reasoning benchmarks, ranging from AIME (competitive math) to MATH and OlympiadBench.

The results consistently show that OPDVR outperforms standard OPD. In the same-architecture setting (distilling a Qwen3-4B model onto itself), OPDVR gains 2.7 points on AIME24 and 2.1 points on AIME25 over standard sampled-token OPD. Remarkably, on AIME24, OPDVR surpasses the teacher model itself. In the cross-architecture setting (distilling a smaller Qwen3-1.7B model from a larger Qwen3-4B teacher), OPDVR delivers gains of 5.5 points on AMC and 1.7 points on MATH500.

The paper also introduces GRPD, which replaces the simple binary reward with a "group-relative advantage" estimate. This variant is even stronger. GRPD outperforms both standard GRPO and standard OPD across the board, with gains as large as 10.9 points on AIME25 over GRPO. This demonstrates that the combination of dense token-level guidance and group-relative correctness checking is a powerful recipe.

Why It Matters: The Big Picture

The takeaways from this paper are significant for anyone building or using advanced AI reasoning models:

  • Simple Integration, No New Hyperparameters: The most immediate practical value is that OPDVR combines the two best post-training techniques without adding complexity. You don't need to tune a weighting factor between "be like the teacher" and "be correct." The ReLU gate handles the alignment automatically. This makes the method easy to adopt in existing training pipelines.
  • The "Teacher-Student-Verifier" Trinity: The paper formalizes a powerful new paradigm. Future methods can rely on this triad: a verifier for truth, a teacher for distribution guidance, and a student for learning. By ensuring the student never learns from the teacher in a way that contradicts the verifier, we get the best of both worlds: the student can surpass the teacher's performance because it's not blindly imitating, and it benefits from the teacher's dense guidance because it's not flying blind with only sparse rewards.
  • A New Failure Mode Fixed: The paper identifies a subtle but important failure mode in standard OPD: it can reinforce the student's overconfidence on wrong answers or punish correct answers if the student is already confident. The ReLU gate explicitly masks out these "conflicting" tokens. This means the model learns more efficiently, focusing its capacity on the steps that actually matter for correctness.
  • What to Watch For: The paper notes that the gating mechanism zeros out about 40-50% of tokens during training. This is a feature, not a bug—it means the model is selectively learning. However, researchers will want to monitor if this gating ever becomes too aggressive, potentially starving the model of necessary distributional guidance. The "inverse-gated" ablation experiment confirms that simply reversing the gate hurts performance, validating that the specific direction of the gate (aligning with verifier correctness) is crucial.

In summary, On-policy Distillation with Verifiable Reward offers a clean, mathematically sound way to marry the sparse but correct signals of reinforcement learning with the dense but unfiltered guidance of distillation. It’s a rare paper that offers a simple fix to a fundamental tension in how we train today's most capable AI models.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →