On-Policy Self-Distillation in Diffusion ModelsExplained for Beginners
Wei Zhou, Xiongwei Zhu, Lingdong Kong +14 more
Abstract
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on-policy self-distillation framework that converts image-level reward guidance into explicit targets for clean-output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fits these targets as detached supervision through finite fitting before an exponential moving average update refreshes the behavior policy. This setup lets us measure target construction and finite realization separately. Controlled same-query experiments show that larger target-construction gains do not necessarily translate into larger realized gains after a single fitting update. Across SD 3.5-M and the step-distilled Z-Image-Turbo, our approach achieves the best final held-out scores in 19 of 20 reward-matched settings across two backbones and ten evaluators. It outperforms the strongest competing method by up to 44.0% and reduces training GPU-hours relative to DiffusionNFT by 40% on SD 3.5-M and 63% on Z-Image-Turbo. These results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.
On-Policy Self-Distillation in Diffusion Models
The Problem
Diffusion models are the workhorses of modern generative AI, producing stunning images from text prompts. But aligning these models with human preferences—making them generate what users actually want rather than just what looks statistically plausible—has proven tricky.
Reinforcement learning offers a pathway, but there's a fundamental mismatch. Reinforcement learning provides a single "reward" signal at the end of the generation process: the generated image is scored, and that score guides improvements. However, diffusion models generate images through a sequence of interdependent denoising steps. The final reward doesn't specify how any intermediate denoising prediction should change. It’s like receiving a grade on a final exam without any feedback on which practice questions you got wrong—you might know your overall score improved, but you can't easily pinpoint which specific step needs adjustment.
Existing methods try to bridge this gap, but they often fall short. Some rely on implicit gradients that are sensitive to how the model is sampled. Others backpropagate reward through only a single late-step prediction, coupling the reward signal too tightly to the optimization process. A method called DiffusionNFT reweights the final outcome, but its target remains an endpoint rather than a concrete instruction for intermediate steps.
The paper identifies a core issue: without explicit intermediate supervision, improving the final outcome requires the model to internally infer how earlier steps should shift. This translation is error-prone and inefficient.
How It Works (The Technical Mechanics)
The paper introduces DiffusionOPSD (On-Policy Self-Distillation), a framework that converts that final image-level reward into explicit, step-by-step targets.
Here is the operational loop, broken into three stages that repeat each training iteration:
1. Trajectory Collection (The "On-Policy" Part)
A frozen "behavior policy" (the current version of the model) generates several image generation trajectories for a set of prompts. For each trajectory, the system records:
- The final generated image (and its reward score).
- A "low-noise query state"—an intermediate point in the denoising process where the image is still mostly noise but semantically meaningful.
- The clean-output prediction the model made at that query state.
Crucially, the model is frozen during this phase. This ensures the data collected reflects the model's current behavior, making the process "on-policy."
2. Target Construction (Turning Reward into Instructions)
For each recorded query state, the system constructs a "positive target" and a "negative target."
- Positive Target: Using the reward gradient (essentially, "how does the reward change if I nudge the image slightly in this direction?"), the system computes a direction in clean-output space that would improve the reward. It takes small steps in that direction, bounded by a trust region (a radius constraint), to create a target prediction that should yield a higher reward.
- Negative Target: Similarly, the system constructs a target that would decrease the reward, acting as a repulsive force.
These targets are computed relative to the anchor—the clean-output prediction the frozen model just made. Because the behavior policy is frozen, the model can freely explore these target directions without worrying about updating its own weights mid-calculation.
3. Finite Fitting (Applying the Lesson)
The trainable policy (the part of the model being updated) now sees these targets as supervised learning signals.
- The model makes a prediction at the same query state.
- It is trained to match the positive target (attraction) and avoid the negative target (repulsion).
- This is "finite fitting"—the model makes a single, bounded update step based on these detached targets.
- Critically, the training process detaches the targets from the reward computation graph. The model isn't being trained to maximize the reward directly; it's being trained to mimic the specific target predictions constructed in step 2.
After this update, the behavior policy is refreshed via an exponential moving average (EMA). The model's weights shift slightly, and in the next outer iteration, new trajectories are collected, new targets are constructed based on the updated model, and the loop repeats.
Key Results & Benchmarks
The paper evaluates DiffusionOPSD on two major backbones: SD 3.5-M (a standard large-scale diffusion model) and Z-Image-Turbo (a distilled, faster model that normally takes 9 steps to generate an image).
Main Results: Setting the State-of-the-Art
Across 20 different reward-evaluation settings (combinations of backbones and evaluators), DiffusionOPSD achieved the best final held-out score in 19 of 20 settings. This is a remarkably consistent result.
- SD 3.5-M: It leads 9 of 10 comparisons. Gains over the strongest competitor range from significant improvements to modest gains. For example, on HPSv3 and VLM-Pairwise, it improves scores by 43.0% and 44.0% respectively.
- Z-Image-Turbo: It leads all 10 comparisons. Gains include 9.7% on Aesthetic, 30.7% on ImageReward, 4.9% on HPSv3, 3.9% on DeQA, and 14.6% on VLM-Pairwise.
Compared to the previous strongest method (DiffusionNFT), DiffusionOPSD reduces the training compute required by 40% on SD 3.5-M and 63% on Z-Image-Turbo to achieve comparable or better results.
Why the "Reversal" Matters
The paper includes a critical nuance regarding how these targets actually translate to improvement. In controlled experiments, they found that a target with a larger "construction gain" (how much it theoretically improves reward in a local calculation) does not always produce a larger "realized gain" after the model actually updates its weights once.
Specifically, on the HPSv2.1 reward, a well-constructed target might actually lead to a worse reward after a single fitting update compared to a random target. This happens on about 62% of prompts. This doesn't invalidate the method; rather, it highlights that target construction and finite model fitting are separate, analyzable stages. The paper uses this separation to diagnose that the reward-gradient direction is the dominant factor in target quality, while the exact noise level of the query state matters much less.
Why It Matters (Key Takeaways)
- Efficiency Gains: By converting image-level rewards into explicit intermediate targets, DiffusionOPSD achieves better alignment with far less compute. The 40-63% reduction in GPU-hours is significant for an industry where training diffusion models is notoriously expensive.
- Diagnosability: The framework's design—separating target construction from finite fitting—allows researchers to pinpoint why a model might not be improving. Is the reward gradient pointing the wrong way? Or is the model failing to execute the update? The paper shows it's usually the gradient direction, making debugging more straightforward.
- Broad Applicability: The method isn't tied to a specific reward model. The jointly trained policy improved all three optimized rewards (PickScore, CLIPScore, HPSv2.1) simultaneously, suggesting a path toward multi-objective alignment without training separate specialists for each metric.
- Limitations to Watch: The method relies on low-noise query states. While the paper argues this is a useful regime (the image is semantically meaningful but the policy still has freedom to change), extremely high noise or extremely conservative target radii can degrade performance. Additionally, the "reversal" phenomenon shows that a single fitting step doesn't always fully realize the constructed target, meaning multiple outer iterations are needed for stable improvement.
In summary, DiffusionOPSD offers a more direct and efficient way to align diffusion models with human preferences. By explicitly constructing "what good looks like" at intermediate steps and teaching the model to hit those marks, it bridges the gap between final rewards and the intermediate denoising process, achieving state-of-the-art results while using fewer computational resources.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →