Systematic Study Compares On-Policy and Off-Policy Learning Distillation Dynamics
fine-tuning reinforcement-learning
| Source: HF Papers | Original article
A new systematic study examines how on-policy versus off-policy learning affects distillation dynamics, addressing gaps in prior comparisons that mixed multiple variables.
A new systematic study has dissected the role of rollout policy in strong‑to‑weak model distillation, contrasting on‑policy and off‑policy learning under tightly controlled conditions. Researchers varied three factors independently—rollout policy, token‑level KL direction, and learning rate—across the Llama 3 and Qwen 2.5 model families while testing on scientific, medical and arithmetic reasoning tasks.
The work challenges the prevailing view that on‑policy learning alone drives the benefits often reported for distillation, such as reduced catastrophic forgetting, sparser updates and better generalisation. In the controlled experiments, on‑policy training did improve performance on harder tasks and appeared to curb incidental “teacher‑style” transfer, but the authors found that learning‑rate settings and the direction of the KL divergence accounted for a substantially larger share of the observed performance differences.
Why this matters: Distillation is a cornerstone of the rapid scaling of large language models, enabling smaller “student” models to inherit capabilities from larger “teacher” systems. Understanding which training knobs truly matter helps practitioners avoid costly trial‑and‑error and could lead to more efficient pipelines that preserve accuracy while limiting forgetting. The findings also temper expectations that simply switching to an on‑policy reinforcement‑learning regime will automatically yield superior students.
What to watch next: The study opens a path for deeper exploration of hyper‑parameter interactions in distillation, especially across diverse model architectures and domains. Industry teams may begin to re‑evaluate their distillation recipes, balancing on‑policy updates against more impactful factors like learning‑rate schedules and KL formulation. As we reported on 1 October in our coverage of UniEvo‑VL, on‑policy self‑distillation is already gaining traction; this latest analysis suggests the next wave will focus on fine‑grained optimisation rather than a blanket policy shift.
Sources
Back to AIPULSEN