RISE: Recursive Improvement via Self-Extrapolating Policy DistillationExplained for Beginners
Yang Li, Semih Yavuz, Shafiq Joty
Abstract
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while self-distillation with privileged conditioning is limited by in-context learning capacity. We propose RISE (Recursive Improvement via Self-Extrapolating Policy Distillation), which constructs a synthetic teacher directly from the model's own RLVR training trajectory. By extrapolating the displacement between the current checkpoint and a trailing anchor---in parameter space or output logit space---RISE converts a sparse outcome-induced parameter update into a dense token-level target, without any external model or privileged conditioning. RISE combines RLVR and OPD in a complementary loop: outcome rewards ground the extrapolation toward correct reasoning, while the extrapolated teacher refines token-level decisions. Moreover, since the teacher is refreshed every iteration as the student improves, distillation becomes a recursive improvement mechanism rather than a one-shot compression step. Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.
Here is a structured explanation of the RISE paper.
1. The Problem
Most language model post-training relies on Reinforcement Learning with Verifiable Rewards (RLVR) to improve reasoning. RLVR provides a sparse signal: the model receives a reward only after generating a complete answer, and that reward is applied uniformly to every token in the response. This creates a "credit assignment" problem—the model cannot easily distinguish which steps of its reasoning were helpful from which were noise.
To make learning denser, researchers use On-Policy Distillation (OPD), which copies the token-level behavior of a "teacher" model. However, RISE identifies a fundamental bottleneck in all existing OPD methods: teacher quality.
- External teachers (stronger, separate models) suffer from distribution mismatch. As the student explores new reasoning paths, the teacher has likely never seen those paths, so its guidance becomes unreliable precisely when the student needs it most.
- Self-distillation (using the model itself as the teacher) often requires "privileged conditioning"—feeding the teacher the correct answer or special metadata. This is limited by the model's in-context learning (ICL) capacity; the model may not be able to effectively incorporate that extra information to produce a helpful token-level signal.
Essentially, the field was accepting a flawed teacher as a given and trying to engineer workarounds for the noise, rather than asking: what should the teacher actually be?
2. How It Works (The Technical Mechanics)
RISE proposes a shift in perspective: the best teacher for a model is its own future self. As the model trains, it moves toward a more capable version of itself. RISE constructs a synthetic teacher by extrapolating the model's own training trajectory.
Here is the core mechanism, explained simply:
The Core Idea: Extrapolation Imagine the model’s parameters taking a step forward during RLVR training. RISE looks at the displacement (the difference) between the current model and a trailing anchor (a previous version of the model). It then amplifies that displacement to project the model to a hypothetical "future" state.
Two Ways to Extrapolate
- Weight-Space (Task Arithmetic): The model literally takes its current parameters and adds a scaled version of the step it just took. If the step was , the "future teacher" has parameters . This is analogous to adding a vector to a coordinate on a map to project forward.
- Logit-Space (Geometric Mixture): Instead of moving parameters, RISE moves the probability distributions. It takes the log-probabilities of the current model and adds a scaled version of the change, creating a "geometric mixture" that amplifies the probability ratios of the newer, better model.
Why This Solves the Teacher Problem Because the teacher is extrapolated from the student's own recent updates, it stays close to the student in parameter and distribution space. It isn't a completely alien model (like an external teacher) nor does it rely on the model's ability to "understand" privileged information (like self-distillation). It is simply the student projected forward along the path it is already carving out.
The Recursive Loop RISE combines this extrapolation with standard RLVR in a two-phase loop:
- RLVR Phase: The model generates rollouts and receives outcome rewards (the sparse signal). This grounds the extrapolation in verified improvement—ensuring the direction points toward correct reasoning.
- OPD Phase: The extrapolated "future teacher" is used to provide dense, per-token supervision. The student mimics the teacher’s token-level decisions.
Critically, because the teacher is refreshed every iteration as the student improves, distillation becomes a recursive improvement mechanism. It isn't a one-time compression step; each round of distillation makes the student better, which provides a fresher, better displacement for the next iteration.
3. Key Results & Benchmarks
RISE was evaluated across mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks. The results show consistent improvements over RLVR-only training and other distillation baselines.
Quantitative Translation The paper reports gains in "Math Avg" accuracy across various model sizes. Here is how to translate those numbers:
- Qwen3-8B: RISE (weight) improved Math Avg from 60.0% to 62.7% (+2.7 percentage points). This means the model answers roughly 2.7% more questions correctly on the standard math benchmark.
- Qwen3-1.7B: RISE (logit) improved Math Avg from 45.4% to 50.2% (+4.8 percentage points).
- OLMo3-7B: RISE (logit) improved AIME'24 from 30.2% to 46.9% (+16.7 percentage points)—a massive jump on a competition benchmark.
Sample Efficiency RISE reaches higher accuracy in fewer training steps. In the early stages of training, where the extrapolated teacher provides a dense signal before RL advantage estimates stabilize, the gap between RISE and GRPO (RLVR-only) is widest. This means RISE gets "more bang for its buck" early on.
Out-of-Distribution (OOD) Preservation Importantly, RISE does not just overfit to the training data. On OOD benchmarks like GPQA and MMLU-Pro, RISE maintains or slightly improves performance (e.g., Qwen3-8B OOD Avg rose from 70.6 to 72.0 with the weight variant). This suggests the token-level refinement sharpens the model without destroying its general capabilities.
4. Why It Matters (Key Takeaways)
- Breaking the Teacher Ceiling: By eliminating the need for external or privileged teachers, RISE removes a hard ceiling on performance. Previous methods using privileged information struggled because the model's in-context learning capacity was the limiting factor. RISE proves that the model's own optimization trajectory contains sufficient structure to serve as a teacher.
- Recursive Self-Improvement: RISE transforms distillation from a static compression step into a loop. As the student gets better, the teacher gets better (because the displacement vector updates). This creates a compounding effect where each iteration potentially yields greater gains than the last.
- Improved Sample Efficiency: Because the extrapolated teacher provides dense token-level supervision from the start, RISE trains faster. It reaches higher accuracy with fewer environment interactions, which is crucial for reducing compute costs.
- Broadened Coverage: The analysis shows RISE doesn't just sharpen the model around answers it already knows; it actually expands the set of problems the model can solve (improving pass@16 more than avg@16 on some benchmarks). This means RISE helps the model tackle harder, novel problems, not just get easier ones right.
What to Watch For The paper notes that the "safe" range for extrapolation () narrows as training converges. If the extrapolation factor is too aggressive late in training, it can overshoot the optimal policy and degrade performance. This is managed via a decaying schedule, but it is a design parameter that requires tuning. Additionally, if the underlying RLVR reward signal is "hackable" (e.g., the model finds shortcuts to get rewards without actual reasoning), extrapolation would amplify those shortcuts.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →