Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressExplained for Beginners
Chen Yang, Haiyuan Wan, Rengrong Xiong +2 more
Abstract
On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher. However, OPD implicitly assumes that teacher-derived rewards are an appropriate proxy for reasoning progress, and therefore treats all teacher feedback equally during policy optimization. While in practice, this assumption does not always hold. We observe that teacher-derived rewards often conflict with genuine reasoning progress, as reasoning steps with clear reasoning advancement may still receive lower distillation rewards, simply due to deviation from teacher's outputs. To address this mismatch, we propose Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward. Distillation rewards are selectively suppressed whenever the two rankings disagree, reducing supervision that conflicts with reasoning progress while preserving effective teacher guidance. Our approach shows consistent improvement over standard OPD especially regarding reasoning performances.
Here is the structured explanation of the research paper "Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress."
1. The Problem
The Gap in On-Policy Distillation (OPD) On-Policy Distillation (OPD) is a technique used to improve language models after they have been initially trained (post-training). The core idea is to pair a student model with a "teacher" model. As the student generates responses, the teacher evaluates each step, rewarding the student for staying close to the teacher’s thought process. This is intended to transfer the teacher’s reasoning capabilities to the student.
However, the paper identifies a critical flaw in this assumption. The authors argue that OPD implicitly treats similarity to the teacher as a proxy for reasoning quality. In practice, this proxy is unreliable.
Why It Matters Outside Academia Think of a student trying to solve a math problem. The teacher might solve it in a specific, rigid way. If the student deviates slightly from the teacher’s exact steps—but takes a clever, valid shortcut that leads to the correct answer—the standard OPD approach might penalize the student. Why? Because the student’s tokens don't match the teacher’s tokens.
Conversely, the student might parrot the teacher’s correct answer without actually understanding the underlying logic. Standard OPD would reward this "imitation" even if the student is just guessing.
For a real-world user, this means a model trained with standard OPD might become a good mimic of the teacher’s style but fail to develop its own robust reasoning skills. It might solve problems by pattern-matching the teacher rather than understanding the logic, or it might give up on valid solution paths simply because they diverge from the teacher’s expectations. The paper argues that this leads to "misleading supervision" during training.
2. How It Works (The Technical Mechanics)
The Core Conflict: Agreement vs. Progress The authors distinguish between two concepts:
- Distillation Compatibility: Does the student’s output match the teacher’s output?
- Reasoning Progress: Does this specific step actually improve the student's chances of getting the right answer?
The Analogy: The GPS Navigation Imagine you are a driver (the student) and you have a GPS (the teacher). Standard OPD tells you: "Stay on the route the GPS suggests." If you take a side street that gets you to the destination faster but isn't on the GPS route, you get "penalized" (reward lowered), even though you are making better progress.
The paper proposes R²-OPD (Reasoning-Progress-Aware Reward Filtering), which acts like a savvy co-pilot. It doesn't just tell you to follow the GPS; it evaluates why you are deviating.
The Mechanics in Three Steps
-
Segmentation (Breaking the Problem Down): The model breaks a long reasoning chain into smaller chunks or "spans." It looks for natural breaks in the text (like sentence boundaries) to identify where one logical thought ends and another begins.
-
Dual Ranking: For each chunk, the system creates two rankings:
- Ranking A (Teacher Signal): How much does this chunk deviate from the teacher's output? (Low deviation = high score).
- Ranking B (Progress Signal): Independently, the model estimates "reasoning progress." It does this by imagining: "If I stop here, how likely am I to eventually get the correct answer?" If this chunk pushes the model closer to the answer, it gets a high progress score.
-
The Filter (Masking): This is the key innovation. The system looks for "conflicts."
-
Scenario: Chunk X has high progress (it moves you closer to the answer) but also high deviation from the teacher (it looks different from the GPS).
-
Action: If Ranking B (Progress) says X is great, but Ranking A (Teacher) says X is bad, the system identifies this as a potential "false positive." It assumes the teacher might be wrong or too rigid here. Consequently, it masks or suppresses the teacher's supervision for that specific chunk.
-
Result: The student is free to follow the logical path (Chunk X) without being penalized for not matching the teacher's exact wording. Meanwhile, if a chunk has high progress and low teacher deviation, the teacher's guidance is preserved.
-
3. Key Results & Benchmarks
The experiments compare the proposed R²-OPD method against standard OPD and several other variants. The primary model tested was DeepSeek-R1-Distill-Qwen-1.5B (the student) trained with JustRL-1.5B (the teacher).
Quantitative Translation
The results are summarized in terms of avg@4 (average accuracy across 4 samples) and pass@4 (probability that at least 1 of 4 samples is correct).
- Standard OPD Baseline: Achieved, for example, 22.50
avg@4and 33.33pass@4on the AIME 2024 benchmark. - R²-OPD Performance: Achieved 32.50
avg@4and 51.83pass@4on the same benchmark.
Plain-Language Impact
- AIME 2024: The improvement is substantial. R²-OPD jumped from 22.50 to 32.50
avg@4. This translates to the model answering approximately 10 more questions correctly out of 100 (or a ~44% relative improvement in accuracy) on this challenging benchmark. - AIME 2025: Similar gains, moving from 20.00 to 25.83
avg@4. - Comparison to other methods: R²-OPD outperformed "Uni-OPD," which was the closest competitor, by about 4.28 points in
avg@4. It also bested standard OPD by 2.51 points inavg@4and 4.46 points inpass@4.
4. Why It Matters (Key Takeaways)
- Higher-Quality Reasoning: By filtering out teacher guidance that conflicts with actual progress, the model is encouraged to explore valid reasoning paths that might deviate from the teacher. This leads to models that are not just imitators but genuine reasoners.
- Robustness to Teacher Flaws: No teacher is perfect. This method makes the student less dependent on the teacher's infallibility. If the teacher makes a mistake or has a quirky style, R²-OPD can essentially "ignore" that part of the teacher's advice if it's hurting the student's ability to solve the problem.
- The Masking Balance: The authors found a "Goldilocks" zone for their filtering. A masking ratio of 30% (suppressing 30% of the most conflicting segments) yielded the best results. Too little masking (10%) left too much conflicting supervision; too much masking (50%) threw out useful teacher guidance, hurting performance.
- General Applicability: While tested on math benchmarks (AIME, OlympiadBench), the mechanism of comparing "imitation" vs. "progress" is domain-agnostic. The authors suggest this framework could be valuable for any reasoning task where the "right way" to solve a problem isn't strictly fixed, such as coding or planning.
What to Watch For: The main limitation noted is the computational cost of estimating "reasoning progress." This requires running multiple "rollouts" (simulated think paths) to calculate the process rewards. As models get larger, this could become expensive. Additionally, the method relies on the quality of the process reward estimator; if the estimator is poor at judging progress, the filtering logic becomes unreliable.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →