Research
Learning from Teacher Continuations at Student States
Learning from Teacher Continuations at Student States (OLIVE) Overview Research area: Natural Language Processing, specifically knowledge distillation for large language model post-training across lon

- arXiv
- 2609.36246
- Published
- 2026-09-28
- Authors
- Haojin Wang, Dylan Zhang, Huaibo Chen, Suhao Yu, Yihang Sun, Zhanyang Jin, Jiaying Ye, Dianqi Li, Prasanna Sattigeri, Kamal Youcef-Toumi, Hao Peng
AI summary
Learning from Teacher Continuations at Student States (OLIVE)Overview
- Research area: Natural Language Processing, specifically knowledge distillation for large language model post-training across long chain-of-thought reasoning and long-horizon agentic tasks.
- Technical level: Advanced. The paper assumes familiarity with supervised fine-tuning, reverse-KL on-policy distillation, covariate shift and compounding errors in behavioral cloning, and DAgger-style learner roll-in with expert rollout.
- Scope: The paper introduces OLIVE (OnLine InterVEntion), an online distillation method in which an evolving student generates a prefix, a teacher continues that prefix, and the student is trained with cross-entropy on only the teacher-generated tokens, plus an asynchronous implementation that overlaps student prefix sampling with teacher generation.
What This Paper Is About
Distillation usually works either by matching a teacher's token-level probability distributions or by training a student on fixed text the teacher generated. Both approaches have problems: fixed teacher trajectories do not match the states the student actually reaches at inference (sequential covariate shift), token-level on-policy distillation grades a student rollout without ever showing what should happen after a correction, and distribution-matching distillation requires access to teacher token probabilities.
OLIVE's goal is to place supervision at the states the student itself reaches, and to refresh that supervision as the student changes, while requiring only teacher-generated text. At each iteration the current student produces a prefix, the teacher autoregressively continues it, and the student is updated with cross-entropy computed solely on the teacher's tokens, with the prompt and student prefix masked out of the loss. The same prefix–continuation split is applied at the level of turns for multi-turn agentic tasks, where the teacher takes over interaction with the environment.
Key Contributions
- OLIVE, an online intervention distillation objective. The paper formulates training where the student rolls out a prefix online, the teacher demonstrates a continuation from that prefix, and the student is trained with a masked cross-entropy loss (Eq. 1) on the teacher text alone; the prompt and student prefix stay in the conditioning context but are excluded from the loss.
- An extension to multi-turn agentic tasks. The prefix–continuation split is applied at the turn level, with
kstudent turns followed by up toMteacher turns executed in the environment, and cross-entropy applied only to the teacher's actions while student turns and all environment observations are masked. - An asynchronous implementation. Student prefix sampling for the next batch overlaps with teacher generation for the current batch, bounded by an asynchronous depth
d(the maximum number of student updates between prefix generation and use of the resulting trace). - Empirical evaluation under matched budgets on RLVE reasoning environments and five AgentGym agentic environments, plus analyses of forgetting on general benchmarks and of plasticity under sequential training.
Main Findings
- Reasoning gains on RLVE (Table 1). With Qwen3-4B-Thinking-2507 as the teacher for all methods, OLIVE raises the Qwen3-1.7B student from a pass@8 of 11.1 to 19.4 (+8.3) and avg@8 from 3.3 to 7.8 (+4.5), compared with OPD at 14.4 (+3.3) pass@8 and 4.6 (+1.3) avg@8 and offline teacher-generated data at 15.0 (+3.9) and 5.6 (+2.3). For the Qwen3-4B student, OLIVE reaches 52.2 (+6.1) pass@8 and 23.5 (+5.4) avg@8, versus OPD at 45.6 (−0.5) and 21.3 (+3.2) and offline data at 53.3 (+7.2) and 19.9 (+1.8). The paper states the largest avg@8 gains among all distillation methods are +4.42 and +5.35 points.
- Agentic task gains (Table 2). With Qwen3-1.7B as student and Qwen3-32B as teacher, OLIVE achieves avg@4 success rates of 40.00 on ALFWorld, 7.50 on ScienceWorld, 55.25 on TextCraft, 67.50 on BabyAI and 39.06 on SearchQA, versus OPD at 22.25, 0.00, 29.50, 43.06 and 29.56. TCoD-B2F, TCoD-F2B and Guided OPD fall between the two on most environments. OLIVE also uses fewer turns in each environment than the other distillation methods.
- ScienceWorld is the largest gain. The paper reports that teacher continuation enables OLIVE to learn where OPD fails, raising ScienceWorld success to 7.5% from a near-zero student that OPD fails to improve, and attributes this to early-turn student mistakes making subsequent turns incorrect under OPD.
- Summary of aggregate improvement. OLIVE lifts hard reasoning tasks by 6% to 8% pass@8 points and agentic benchmarks by 7% to 22% avg@4 gains, while introducing only a 0.9% average performance drop on general benchmarks.
- Flat top-K overlap under OPD (Figure 3). For the Qwen3-1.7B student and Qwen3-4B-Thinking-2507 teacher, the top-K overlap ratio on the validation set stays nearly flat during OPD training, moving from 0.707 to 0.713, which the paper links to different thinking behaviors between student and teacher.
- Efficiency. Asynchronous OLIVE reduces total training time by 23.8% relative to synchronous OLIVE and by 28% compared with OPD, with minimal performance degradation. OLIVE matches OPD's training efficiency while improving performance, because it does not require the teacher to complete the reasoning process or compute reverse KL on all student tokens.
- Less forgetting, more improvement (Section 5.1). Compared against an offline OEC-style variant and an offline SFT baseline on RLVE, OLIVE achieves better task performance and less forgetting on four general benchmarks: AIME25 (math, avg@16), LiveCoding Bench v6 (code, avg@8), IF-Eval (instruction following) and GPQA Diamond (science).
- Continued improvement online (Section 5.2). Using only text from GPT-5.4-mini as teacher with Qwen3-1.7B as student on ScienceWorld, offline distillation plateaus after 2 epochs while OLIVE keeps improving and surpasses offline SFT from the same teacher by 13% after 5 epochs.
- Plasticity under sequential training (Figure 8). Training sequentially on each agentic environment for 5 epochs with Qwen3-4B, the gap is largest on SciWorld, the last environment in the sequence, where OLIVE improves over offline distillation by 13.0 points, indicating that offline trajectories increasingly mismatch the states the shifted policy visits.
Methodology in Plain English
OLIVE works in a repeating loop. First, the current student model is given a prompt and writes the beginning of an answer: a fixed-length prefix of k tokens for reasoning tasks, or k interaction turns for agentic tasks. The prefix is sampled from the student's own policy, so it reflects where the student currently is, and the sampled prefixes are kept without filtering.
Second, the teacher model receives the prompt plus the student's prefix and writes a continuation of at most M tokens (or M turns). Because the teacher conditions each new token on its own preceding choices, the continuation shows how to proceed after a correction, not just that a correction is needed. The paper deliberately uses partial continuations rather than full verified solutions, to bound the cost of online teacher generation.
Third, the teacher's continuation is tokenized with the student's tokenizer and appended to the prompt and student prefix. The student is trained with cross-entropy only on the teacher's tokens; the prompt and prefix remain as context but are masked out of the loss, and the sampled sequences are held fixed during the update. Because only teacher-generated text is needed, no teacher token probabilities or logits are required, and teacher and student tokenizers may differ — this enables black-box distillation from API models.
For agentic tasks, the student acts in the environment for k turns, producing a history of actions and observations, and then the teacher takes over for up to M turns, with each teacher action executed in the environment to produce the next observation. During the update, the full interaction history is context, student turns and environment observations are masked, and cross-entropy applies only to teacher actions. The paper finds student prefixes of k = 5 or 10 turns depending on the environment, followed by up to M = 5 teacher turns, work well.
Because the student policy changes at every update, prefixes are regenerated throughout training, following the DAgger principle of supervising at learner-visited states. To cut idle time, the asynchronous implementation runs student prefix sampling for the next batch while the teacher generates continuations for the current batch; the maximum staleness is bounded by depth d (set to 3 in the reasoning experiments).
Experimental setup: RLVE provides 18 reasoning environments with 500 training problems each, giving a pool of 9K hard problems, with 10 same-difficulty test problems per task. Students are Qwen3-1.7B and Qwen3-4B with thinking enabled, and the reasoning teacher is Qwen3-4B-Thinking-2507. For a matched budget across methods, distilled tokens per rollout are capped at 7168; OLIVE uses prefix length 4096 and continuation length 1024 by default. Agentic experiments use five AgentGym environments (ALFWorld, ScienceWorld, TextCraft, BabyAI, SearchQA), with SearchQA given a 6K-problem training set and 400 held-out evaluation problems, Qwen3-1.7B as student and Qwen3-32B as teacher, maximum turns of 30 for ALFWorld, TextCraft and ScienceWorld, 20 for BabyAI and 16 for SearchQA, and ReAct-style action generation per turn.
Why This Matters
- Impact on research. The paper argues that where supervision is placed, and whether it is refreshed as the student changes, is an important axis of distillation design alongside the form the supervision takes. It connects LLM distillation to classical imitation-learning ideas about learner roll-in with expert rollout and to expert-intervention methods for language-model agents, and it offers a text-only alternative for settings where teacher logits are unavailable.
- Real-world applications:
- Distilling from closed API teachers (such as GPT-5.4-mini in the paper's study) that expose only generated text, since OLIVE needs no token probabilities.
- Post-training agents for multi-turn interactive environments such as ALFWorld, ScienceWorld, TextCraft, BabyAI and SearchQA, where early student errors otherwise derail whole episodes.
- Training reasoning models on hard synthetic problems with tunable difficulty, as in the RLVE setup, where teacher trajectories alone under-supervise student-specific failure states.
- Sequentially adapting a single student across multiple task environments while limiting loss of plasticity and forgetting of general capabilities.
- Industry relevance. The asynchronous implementation makes online teacher generation practical: it reduces training time by 23.8% over synchronous OLIVE and 28% over OPD, and the paper reports comparable GPU-hour cost to OPD with better reasoning performance on 8 H200 GPUs. Black-box compatibility also removes the requirement to host or expose teacher logits, which matters when the strongest available teacher is an API model.
Future Directions
- The paper is marked "Ongoing work," and the full range of open problems is not enumerated; the published content does not report a formal limitations section.
- The asynchronous depth
dtrades overlap against mismatch between the policy that generated a prefix and the policy being trained; the paper usesd= 3 for reasoning and evaluates the efficiency–performance tradeoff, leaving how far this can be pushed as an open question. - How OLIVE scales beyond the student and teacher sizes studied here (Qwen3-1.7B and Qwen3-4B students, Qwen3-4B-Thinking-2507, Qwen3-32B and GPT-5.4-mini teachers) is not reported.
- Choosing prefix length
kand continuation budgetMis currently done per environment (5 or 10 student turns, 5 teacher turns; 4096 prefix tokens and 1024 continuation tokens for reasoning); the paper does not report a general procedure for setting these.
Target Audience
Readers who benefit most are researchers and engineers working on LLM post-training, distillation and reinforcement-learning-style fine-tuning, particularly those training reasoning models or multi-turn agents. It is also relevant to practitioners who can only access a teacher through an API that returns text, and to those studying forgetting, plasticity and continual adaptation of language models. The paper is written at an advanced level and assumes background in supervised fine-tuning, on-policy distillation and imitation learning.
Authors’ abstract
We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE's total training time by 23.8\%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13\% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.