+15.8%
WebShop score over OPD with a 1.5B student (66.3 → 76.8)
arXiv preprint · 2026
1University of California, Irvine 2University of Michigan, Ann Arbor *Equal contribution
+15.8%
WebShop score over OPD with a 1.5B student (66.3 → 76.8)
90.0 / 88.8
ALFWorld seen / unseen success (%) with a 3B student, best among distillation methods
+2.7 to +8.0
exact-match points over OPD on all four multi-hop QA datasets
5 of 6
ALFWorld/WebShop settings where UOPD needs the fewest interaction turns among distillation methods
Abstract
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation.
We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback–Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to 15.8% relative to standard OPD.
Method in motion
Both students propose the same actions until turn 3, where the student heads to the microwave although the task asks for a clean mug. OPD executes that action and only adjusts probabilities afterwards; UOPD sees the teacher’s high uncertainty, executes the teacher’s action instead, and trains the student to imitate it.
Illustrative ALFWorld-style episode with schematic uncertainty scores, built to show the mechanism. Apart from this episode, the intervention-rule explorer, and the Proposition 1 construction, every number on this page is reported in the paper. Download this episode as a video (WebM, 34 s).
Method
UOPD keeps standard OPD on every turn the teacher is comfortable with and intervenes only where the teacher is uncertain about the student’s proposal. One decision sets both the executed action and the learning objective (Figure 1).
The student samples \(a^S_t \sim \pi_\theta(\cdot\mid h_t)\). On ALFWorld and WebShop the frozen teacher scores it with its length-normalized negative log-likelihood, \(\delta_t = -\tfrac{1}{|a^S_t|}\log \pi_T(a^S_t\mid h_t)\); Search uses the student–teacher confidence gap.
The threshold is a quantile of recent scores, \(\tau_n=\mathrm{Quantile}_{1-r_n}(\mathcal D_n)\), so the scheduled rate \(r_n\) sets how often the teacher steps in while uncertainty sets where.
If \(\delta_t>\tau_n\), UOPD executes a teacher action \(a^T_t\sim\pi_T\) and trains the student on it with SFT; otherwise it executes \(a^S_t\) and applies the OPD reverse-KL term.
Per-turn objective
Scheduled intervention rate
Why SFT on teacher actions
Pick a schedule and move through training. The threshold is recomputed from the score buffer, and every turn above it is corrected by the teacher.
Motivation
A controlled ALFWorld study keeps both models fixed (Qwen2.5-3B student, GiGPO-Qwen2.5-7B teacher) and replaces exactly one student action with a teacher action, either at the first low-confidence step or at a random step, on the same 93 tasks.
Results
RL-trained GiGPO-Qwen2.5-7B-Instruct teachers are distilled into Qwen2.5-3B and Qwen2.5-1.5B students and compared with OPD, both TCOD curricula, and FTB-OPD. Values are means ± standard deviations over three evaluation seeds (paper Table 1).
A Qwen2.5-1.5B student distilled from SearchR1-Qwen2.5-7B; exact match on Search-R1’s four multi-hop test sets (22,523 questions).
Wall-clock hours for the same number of steps on eight RTX PRO 6000 GPUs with 3B students. UOPD decides before acting, so it never pays for paired continuations.
Ablations
Qwen2.5-3B students. The default uses teacher confidence, \(\beta=1\), teacher execution, and a 30% → 5% schedule; each variant changes one component (paper Table 2).
At the same scheduled rate, uncertainty-selected turns beat random turns: unseen ALFWorld 88.8 vs. 83.8, WebShop score 82.4 vs. 77.3 with 6.5 vs. 7.4 turns.
Dropping the SFT term (\(\beta=0\)) costs 8.0 points on unseen ALFWorld; teacher execution alone does not recover the benefit, and \(\beta=2\) adds nothing.
Executing the teacher’s action lifts the WebShop score from 77.6 to 82.4 and trims turns from 7.2 to 6.5, even though success rates stay similar.
A 50% → 30% schedule slightly raises seen ALFWorld but hurts unseen ALFWorld and WebShop; a constant 15% stays competitive on ALFWorld.
ALFWorld, 3B student, 30% → 5% over 120 rollout batches. Corrections become rarer while their position stays variable: UOPD keeps choosing turns by uncertainty.
Theory
A correction is both an action to execute and a target to imitate. The paper analyzes each role.
Theorem 1 · execution
For any uncertainty intervention rule and teacher policy in a finite-horizon environment, the assisted policy \(\pi_G\) satisfies
If teacher actions have an average advantage of at least \(\gamma>0\) at the selected turns, then \(J(\pi_G)-J(\pi_S)\ge\gamma\,\mathbb E_{\pi_G}[N_{\mathrm{int}}]\ge 0\): on average the gain is at least \(\gamma\) per intervention, before any student update. Not every correction has to help.
Proposition 1 · supervision
There is a one-step, three-action problem with full-support policies and a teacher better than the student where one small gradient step of reverse KL lowers return while one step of teacher-action SFT raises it:
Explore the exact construction from the proof below.
Actions \(a\) and \(b\) succeed, \(c\) fails. The student starts uniform; the teacher is \(\pi_T=(v(1-\epsilon),\,v\epsilon,\,1-v)\). Every number below is computed exactly from the paper’s gradients \(G_{\mathrm{OPD},i}=p_i(\ell_i-\bar\ell)\) and \(G_{\mathrm{SFT},i}=p_i-u_i\).
Citation
@article{zhang2026uopd,
title = {{UOPD}: Uncertainty-Aware Intervention for On-Policy
Distillation of Multi-Turn Agents},
author = {Zhang, Wenbo and Xu, Pengcheng and Du, Weizhi
and Zhang, Jing and Cai, Hengrui},
journal = {arXiv preprint arXiv:2609.34036},
year = {2026}
}