UOPD

arXiv preprint · 2026

UOPD Uncertainty-Aware Intervention for On‑Policy Distillation of Multi‑Turn Agents

  • Wenbo Zhang1,*
  • Pengcheng Xu1,*
  • Weizhi Du2
  • Jing Zhang1
  • Hengrui Cai1

1University of California, Irvine 2University of Michigan, Ann Arbor *Equal contribution

Figure 1: OPD versus UOPD

Figure 1. OPD versus UOPD. OPD executes student actions \(a^S_t\) given observations \(o_t\) and distills with reverse KL. UOPD uses the teacher’s uncertainty on each proposal and an adaptive threshold to decide when the teacher intervenes; the decision sets both the executed action and the objective: reverse KL for student actions, or SFT on teacher actions \(a^T_t\), which minimizes forward KL in expectation. In the interactive view, drag τ to change which turns are corrected and hover any component for details (per-turn scores are illustrative).

Highlights

+15.8%

WebShop score over OPD with a 1.5B student (66.3 → 76.8)

90.0 / 88.8

ALFWorld seen / unseen success (%) with a 3B student, best among distillation methods

+2.7 to +8.0

exact-match points over OPD on all four multi-hop QA datasets

5 of 6

ALFWorld/WebShop settings where UOPD needs the fewest interaction turns among distillation methods

Abstract

Correct the student where it matters

On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation.

We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback–Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to 15.8% relative to standard OPD.

Method in motion

One episode, two distillation strategies

Both students propose the same actions until turn 3, where the student heads to the microwave although the task asks for a clean mug. OPD executes that action and only adjusts probabilities afterwards; UOPD sees the teacher’s high uncertainty, executes the teacher’s action instead, and trains the student to imitate it.

Illustrative ALFWorld-style episode with schematic uncertainty scores, built to show the mechanism. Apart from this episode, the intervention-rule explorer, and the Proposition 1 construction, every number on this page is reported in the paper. Download this episode as a video (WebM, 34 s).

Method

Uncertainty decides who acts and what is learned

UOPD keeps standard OPD on every turn the teacher is comfortable with and intervenes only where the teacher is uncertain about the student’s proposal. One decision sets both the executed action and the learning objective (Figure 1).

1

Score the proposal

The student samples \(a^S_t \sim \pi_\theta(\cdot\mid h_t)\). On ALFWorld and WebShop the frozen teacher scores it with its length-normalized negative log-likelihood, \(\delta_t = -\tfrac{1}{|a^S_t|}\log \pi_T(a^S_t\mid h_t)\); Search uses the student–teacher confidence gap.

2

Decide with a quantile

The threshold is a quantile of recent scores, \(\tau_n=\mathrm{Quantile}_{1-r_n}(\mathcal D_n)\), so the scheduled rate \(r_n\) sets how often the teacher steps in while uncertainty sets where.

3

Act and learn accordingly

If \(\delta_t>\tau_n\), UOPD executes a teacher action \(a^T_t\sim\pi_T\) and trains the student on it with SFT; otherwise it executes \(a^S_t\) and applies the OPD reverse-KL term.

Per-turn objective

$$\ell^{\mathrm{UOPD}}_t=\begin{cases}\log\pi_\theta(a^S_t\mid h_t)-\log\pi_T(a^S_t\mid h_t), & \delta_t\le\tau_n \quad\text{(execute student, reverse KL)}\\[4pt] -\beta\,\log\pi_\theta(a^T_t\mid h_t), & \delta_t>\tau_n \quad\text{(execute teacher, SFT)}\end{cases}$$

Scheduled intervention rate

$$r_n=r_{\mathrm{start}}+(r_{\mathrm{end}}-r_{\mathrm{start}})\min\!\Big(\frac{n}{N_{\mathrm{decay}}},1\Big)$$

Why SFT on teacher actions

$$\mathbb E_{a^T\sim\pi_T}\!\big[-\log\pi_\theta(a^T\mid h)\big]=\mathcal H(\pi_T)+\mathrm{KL}(\pi_T\,\|\,\pi_\theta)$$

Try the intervention rule

Pick a schedule and move through training. The threshold is recomputed from the score buffer, and every turn above it is corrected by the teacher.

Motivation

One correction at a low-confidence step changes the rest of the rollout

A controlled ALFWorld study keeps both models fixed (Qwen2.5-3B student, GiGPO-Qwen2.5-7B teacher) and replaces exactly one student action with a teacher action, either at the first low-confidence step or at a random step, on the same 93 tasks.

  • Uncertainty varies within an episode. Teacher confidence falls and recovers along student rollouts, so some steps are far less supported than others.
  • One correction helps later actions too. After a single intervention the teacher is more confident in the student’s own subsequent actions, with no parameter update.
  • Timing matters. With the same amount of teacher help, intervening at a low-confidence step nearly doubles success (13.98% → 25.81%), while a random step barely moves it (16.13%).
Figure 2. (a) Teacher confidence fluctuates across turns of student-only rollouts. (b) Replacing one low-confidence action raises teacher confidence on the student’s subsequent actions. (c, d) The intervention increases task success and reduces backtracking compared with student-only rollouts and random intervention.

Results

Better agents, shorter episodes, at two student sizes

RL-trained GiGPO-Qwen2.5-7B-Instruct teachers are distilled into Qwen2.5-3B and Qwen2.5-1.5B students and compared with OPD, both TCOD curricula, and FTB-OPD. Values are means ± standard deviations over three evaluation seeds (paper Table 1).

Multi-hop search

A Qwen2.5-1.5B student distilled from SearchR1-Qwen2.5-7B; exact match on Search-R1’s four multi-hop test sets (22,523 questions).

Training cost

Wall-clock hours for the same number of steps on eight RTX PRO 6000 GPUs with 3B students. UOPD decides before acting, so it never pays for paired continuations.

Figure 4. Training dynamics with 3B students. UOPD improves rollout quality early in training and its trajectory-level teacher–student KL falls faster than OPD and TCOD-B2F. B2F starts high thanks to teacher prefixes but drops as the prefixes shorten.

Ablations

What each component contributes

Qwen2.5-3B students. The default uses teacher confidence, \(\beta=1\), teacher execution, and a 30% → 5% schedule; each variant changes one component (paper Table 2).

Where to intervene

At the same scheduled rate, uncertainty-selected turns beat random turns: unseen ALFWorld 88.8 vs. 83.8, WebShop score 82.4 vs. 77.3 with 6.5 vs. 7.4 turns.

Learn from corrections

Dropping the SFT term (\(\beta=0\)) costs 8.0 points on unseen ALFWorld; teacher execution alone does not recover the benefit, and \(\beta=2\) adds nothing.

Execute corrections

Executing the teacher’s action lifts the WebShop score from 77.6 to 82.4 and trims turns from 7.2 to 6.5, even though success rates stay similar.

Do not over-intervene

A 50% → 30% schedule slightly raises seen ALFWorld but hurts unseen ALFWorld and WebShop; a constant 15% stays competitive on ALFWorld.

Intervention dynamics during training

ALFWorld, 3B student, 30% → 5% over 120 rollout batches. Corrections become rarer while their position stays variable: UOPD keeps choosing turns by uncertainty.

Figure 7. Left: actual episode-averaged intervention rate and the scheduled target. Right: mean intervention turn among corrected trajectories.

Theory

Two roles of a teacher correction

A correction is both an action to execute and a target to imitate. The paper analyzes each role.

Theorem 1 · execution

Selective intervention improves rollout return

For any uncertainty intervention rule and teacher policy in a finite-horizon environment, the assisted policy \(\pi_G\) satisfies

$$J(\pi_G)-J(\pi_S)=\mathbb E_{\pi_G}\Big[\sum_t I_t\big(Q^{\pi_S}_t(h_t,a^T_t)-Q^{\pi_S}_t(h_t,a^S_t)\big)\Big].$$

If teacher actions have an average advantage of at least \(\gamma>0\) at the selected turns, then \(J(\pi_G)-J(\pi_S)\ge\gamma\,\mathbb E_{\pi_G}[N_{\mathrm{int}}]\ge 0\): on average the gain is at least \(\gamma\) per intervention, before any student update. Not every correction has to help.

Proposition 1 · supervision

Imitating the teacher can help where reverse KL hurts

There is a one-step, three-action problem with full-support policies and a teacher better than the student where one small gradient step of reverse KL lowers return while one step of teacher-action SFT raises it:

$$J_{t,h}(\pi^+_{\mathrm{OPD}})<J_{t,h}(\pi_S)<J_{t,h}(\pi^+_{\mathrm{SFT}}).$$

Explore the exact construction from the proof below.

Proposition 1, live

Actions \(a\) and \(b\) succeed, \(c\) fails. The student starts uniform; the teacher is \(\pi_T=(v(1-\epsilon),\,v\epsilon,\,1-v)\). Every number below is computed exactly from the paper’s gradients \(G_{\mathrm{OPD},i}=p_i(\ell_i-\bar\ell)\) and \(G_{\mathrm{SFT},i}=p_i-u_i\).

Figure 6. Mode seeking and mode covering. Reverse KL refines the student around a teacher-supported mode; forward KL, applied through SFT on corrections, covers teacher behaviors the student rarely generates (schematic).

Paper

Read the full paper

First page of the UOPD paper

UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents

Wenbo Zhang*, Pengcheng Xu*, Weizhi Du, Jing Zhang, Hengrui Cai

arXiv:2609.34036 · cs.LG · 2026

Main text, pages 1–12

Citation

BibTeX

@article{zhang2026uopd,
  title   = {{UOPD}: Uncertainty-Aware Intervention for On-Policy
             Distillation of Multi-Turn Agents},
  author  = {Zhang, Wenbo and Xu, Pengcheng and Du, Weizhi
             and Zhang, Jing and Cai, Hengrui},
  journal = {arXiv preprint arXiv:2609.34036},
  year    = {2026}
}