Research
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
Overview Research area: Reinforcement learning and self-distillation for LLM agents, specifically training a small "advisor" model that steers a frozen, API-served "executor" model through natural-lan

- arXiv
- 2609.38142
- Published
- 2026-09-29
- Authors
- Rishabh Agrawal, Hejie Cui, Shasha Li, Shanchan Wu, Sercan Ö. Arık
AI summary
Overview
Research area: Reinforcement learning and self-distillation for LLM agents, specifically training a small "advisor" model that steers a frozen, API-served "executor" model through natural-language advice.
Technical level: Intermediate. The paper combines a formal analysis (a lemma and two theorems on a simplified shared-parameter model) with a concrete training pipeline, so readers should be comfortable with reinforcement learning and distillation terminology, but the empirical protocol and the core idea are described accessibly.
Scope: This paper proposes Advisor Self-Distillation (AdviSD), a method that trains a Qwen3-8B advisor to steer frozen frontier executors by combining outcome-based GRPO with selectively targeted self-distillation from a feedback-conditioned copy of the advisor, and evaluates it on BFCL-v3 and EnvScaler with Gemini 3.7 Flash and Claude Sonnet 4.6 executors.
What This Paper Is About
Frontier language models are typically served through APIs, so their weights cannot be changed; when such a model acts as an agent, a smaller trainable advisor can adapt it from the outside by recommending what the frozen agent should do next. The problem is that a plausible correction proposed by reflection need not actually change what the executor does, yet learning from such a correction can still shift the advisor's shared parameters and affect its advice in other contexts. The paper asks which corrections are worth learning from for a given advisor–executor pair, and answers with a selection rule based on how much the issued advice changes the advisor's prediction of the executor's recorded response.
Key Contributions
-
A theoretical account of correction selection. Lemma 1 separates agreeing with a feedback-conditioned teacher from improving execution, showing that reverse-KL descent toward the actual teacher follows the value-ascent direction plus a residual term that can reinforce or oppose it. Theorems 1–2 then analyze repeated learning in a shared-parameter model with sensitive and insensitive situations.
-
A targeted selection mechanism. AdviSD scores the same recorded executor response under two contexts that differ only in whether the issued advice is present, computing a per-decision contrast (Equation 3), and uses the magnitude of that contrast against a donor-calibrated threshold to choose which reflection-flagged decisions receive supervision.
-
Self-distillation that complements outcome-based RL. A feedback-conditioned copy of the pre-update advisor teaches a student that sees only the original context, with GRPO advantages left unchanged on all episodes. Selection requires neither executor likelihoods nor additional executor rollouts.
-
Empirical validation including a matched-count random control. Experiments with Qwen3-8B advisors for Gemini and Claude on BFCL-v3 and EnvScaler, plus out-of-domain and transfer evaluations, test whether choosing particular corrections helps beyond merely reducing how often supervision occurs.
Main Findings
-
AdviSD leads in-domain. With Qwen3-8B advisors, AdviSD has the highest BFCL-v3 average and EnvScaler score for both executors (Table 1). It exceeds advisor-GRPO by 4.2–6.4 percentage points on BFCL-v3 and by 3.9–5.1 score points on EnvScaler.
-
Gains are larger in BFCL's augmented categories. For both executors, the improvement over GRPO is larger in the augmented categories than in Base, suggesting greater benefit when tools are missing (Missing Function), required information is missing (Missing Parameter), or contexts are longer (Long Context).
-
Untrained small-advisor advice can hurt. Untrained Qwen3-8B advice lowers all four aggregate means relative to standalone execution, while frozen frontier advisors raise them.
-
A trained small advisor can beat frontier advisors. After AdviSD training, the smaller Qwen3-8B advisor outperforms the frontier advisors in every comparison in Table 1 except Missing Parameter with the Claude executor.
-
Shared trajectory-level feedback trails GRPO. Adapted SDPO and DistIL reuse the same successful-rollout feedback to supervise the advisor's generated tokens across turns, yet both trail GRPO in aggregate, which the authors say motivates decision-level targeting.
-
Selection beats reduced supervision at matched counts. Across settings, AdviSD shows a 2.5–4.9-point advantage over matched-count random selection, supporting the value of the selection rule rather than the mere reduction of supervision.
-
Theory predicts the selection effect. In the two-teacher model, decreasing the retention ratio r_I/r_S strictly increases both the equilibrium θ* and expected success J(θ*); changing the episode size, the proposal cap, or scaling both retention probabilities by the same admissible positive factor leaves the limit unchanged. With reward learning added, selective retention gives higher equilibrium θ and J than no gating.
-
Random thinning is not equivalent. Corollary 1 notes that retaining every proposal independently with the same fixed probability preserves the teacher mixture and the distillation-only limit, yet can improve eventual performance when reward learning is present — which is why the paper includes a matched-count random control.
-
Out-of-domain transfer without retraining. BFCL-selected systems improve out-of-domain macro-averages by 2.7–3.6 points over standalone execution, and AdviSD transfers across executor versions and model families, exceeding transferred GRPO by 3.1 percentage points in both cross-family directions.
-
Ablations isolate the gate and the abstention bypass. In Table 2, "no gate" and "matched-count random" sit near GRPO, the "inverted gate" is at or below GRPO, and removing the abstention bypass reduces performance relative to full AdviSD (for example, 66.1 versus 67.8 on BFCL-v3 with the Gemini executor, and 83.3 versus 84.8 on EnvScaler with the Gemini executor).
-
Note on the out-of-domain table. Table 3 in the provided content is truncated; the complete per-benchmark numbers for ACEBench, ToolHop, τ²-bench, and RoTBench are not fully reported in the available text beyond the partial rows shown.
Methodology in Plain English
The setup has two models: a frozen executor that actually performs the task, and a trainable advisor that decides before each executor response whether to issue advice or abstain with a special <NO_ADVICE> token. Advice is placed in a temporary request rather than the executor's persistent history, and the advisor makes a new decision before every response, including text-only responses and those following tool results. The advisor is trained with GRPO on the episode reward, where the advantage for an episode i is A_i = (R_i − mean R) / (std R + ε_A) applied to every advisor token in that episode.
AdviSD adds a second learning signal that is selective. First, a reflector inspects imperfect episodes and flags at most b_refl = 5 advice decisions with correction feedback. Next, a scoring step asks: how much did the advice I actually issued change my prediction of what the executor did? The pre-update advisor scores the recorded executor response token by token under two contexts — one containing the issued advice, one without it — and averages the log-probability difference across tokens. The magnitude of this contrast is used as a signal for whether the decision is worth supervising.
To set the gate, the authors run pilot rollouts with the initial advisor and, at decisions where advice was issued, replace it with advice from another task (donor advice). Donor advice is scored but never sent to the executor. The threshold is an empirical quantile of the donor contrast magnitudes, with u_d = 0.95, and it stays fixed during training. Two kinds of flagged decisions are retained: original abstentions (which bypass scoring because the with-advice and without-advice contexts are identical when no advice was issued, so the contrast is zero), and decisions whose contrast exceeds the threshold.
At each retained decision, a teacher — a gradient-stopped copy of the pre-update advisor — sees the original context with feedback prepended, while the student sees only the original context. Both predict along the originally sampled advice, and the student is trained to match the teacher over the pre-update student's top-K tokens with a temperature T_SD. The auxiliary loss is averaged within each decision, then within each episode, then across episodes, and added to the GRPO loss with weight λ_s. Reflection and scoring are used only during training; at deployment the advisor simply decides before each executor turn whether to advise or abstain.
Evaluation uses three independent training runs per trained method; each run's validation-selected checkpoint is evaluated four times on the test set, and the reported mean ± sample SD is over the three run averages. GEPA is aggregated the same way using three independent prompt searches, and the no-advisor and frozen-advisor controls use three groups of four evaluations of the same system.
Why This Matters
The work addresses a practical constraint: the strongest models are often accessible only through APIs, so any adaptation must happen outside the model's weights. It also raises a conceptual point that applies broadly to learning from feedback — that a correction which is plausible, or which a feedback-conditioned teacher endorses, is not necessarily a correction that improves task execution, and that teaching a model on all such corrections can degrade its behavior elsewhere through shared parameters.
Real-world applications:
- Tool-using enterprise agents. The BFCL-v3 categories involving missing functions and missing parameters model exactly the situation where an agent must ask for or infer missing information, common in API-driven workflows.
- Customer-facing booking and scheduling agents. The paper's own motivating example is a travel booking scenario where an executor guesses a passenger identifier instead of looking it up; targeted advice fixes the specific failure rather than restating procedures the executor already follows.
- Model-agnostic agent adaptation. Because advising works through a response-level interface and requires no executor likelihoods or extra executor rollouts, the same trained advisor can be applied to several executors, including across model families.
- Improving weaker executors without retraining. A small trained advisor can raise the performance of a frozen executor it was not trained against, which matters where fine-tuning is impossible or prohibitively costly.
Industry relevance: the paper targets exactly the deployment pattern of frontier models behind APIs, where organizations control the surrounding orchestration rather than the model weights. The claimed transfer across executor versions and families is directly relevant to teams that switch model providers or update model versions and want their agent scaffolding to keep working.
Future Directions
- Replacing the calibrated gate with a learned selector. The current threshold is an empirical donor quantile fixed before training; whether the selection rule can be learned or adapted online is not reported.
- Extending the theory past the two-teacher model. The analysis in Sections 4 and 6 assumes fixed teacher preferences, fixed executor behavior under any given advice, two advice options, and fixed success probabilities in (0, 1); the appendices provide scope and extensions, but richer settings remain open.
- Understanding exposure effects during training. Section 6 notes that a higher learning limit need not mean faster learning from the start, since below the distillation-only equilibrium reducing supervision can slow early progress; how to schedule supervision to capture both effects is not resolved.
- Broader transfer evaluation. The paper reports transfer across executor versions and model families; whether the approach transfers to executors with substantially different tool interfaces, or to advisors other than Qwen3-8B, is not reported.
Target Audience
Researchers and engineers working on LLM agents, reinforcement learning from task rewards, and feedback-based self-distillation will benefit most, particularly those interested in improving frozen models they cannot fine-tune. The paper is also relevant to practitioners deploying tool-using agents behind APIs who need to port agent scaffolding across model versions and vendors. Readers primarily interested in clean empirical results without formal analysis should note that Sections 4 and 6 are theory-heavy, though the method description in Sections 5 and 7 stands on its own.
Authors’ abstract
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.