Research
PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation
Overview Research area: Machine learning / natural language processing — knowledge distillation and reinforcement learning for small language models on vertical-domain few-shot intent classification.
- arXiv
- 2610.11167
- Published
- 2026-10-08
- Authors
- Heng Li, Yong Zhang, Ning Cheng, Zhigen Li, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao
AI summary
Overview
Research area: Machine learning / natural language processing — knowledge distillation and reinforcement learning for small language models on vertical-domain few-shot intent classification.
Technical level: Advanced. The paper assumes familiarity with on-policy distillation, KL divergence, GRPO-style policy optimization, and reward-normalized advantage estimation.
Scope: The paper proposes and empirically evaluates PIVOT, a scheduling framework that decides per sample when to stop distilling from a teacher and start refining with GRPO, using teacher-evaluated sequence perplexity as the routing signal, evaluated on Banking77 and HWU64 with Qwen2.5-class models.
What This Paper Is About
Small language models adapted to professional domains such as finance struggle when only a handful of labeled examples per class are available. Existing pipelines typically distill from a large teacher model first and then switch to reinforcement learning, but they make that switch globally and at a fixed point for every sample. PIVOT asks instead whether different samples should switch at different times, and uses the teacher's own perplexity on the student's generated responses to decide.
Key Contributions
- The paper shows that prolonged teacher-guided distillation does not necessarily produce better downstream GRPO refinement under vertical-domain few-shot adaptation.
- It proposes PIVOT (Perplexity-Informed Transition Optimization), a progressive sample-level KD-to-RL transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity.
- It reports that on Banking77 and HWU64, PIVOT outperforms continued OPD and a globally synchronized OPD-to-GRPO baseline under the same number of post-warm-up student optimization steps, with more stable optimization dynamics.
- It provides ablation evidence on routing rules and transition schedules, plus a teacher-generated data augmentation comparison to argue the gains are not explained by additional synthetic supervision alone.
Main Findings
- Main accuracy results (Table 1): Under 480 post-warm-up student optimization steps, PIVOT reaches 80.26 on Banking77 and 85.04 on HWU64, versus 78.05 and 83.09 for globally synchronized OPD-to-GRPO, and 72.47 and 79.55 for continued OPD. Teacher zero-shot inference scores 73.25 on Banking77 and 79.65 on HWU64, so PIVOT exceeds the teacher on both datasets.
- Other baselines: SeqKD scores 58.70 (Banking77) and 70.17 (HWU64); GRPO alone scores 67.66 and 77.60.
- Scaling analysis (Table 2): PIVOT improves over OPD-to-GRPO in every configuration tested: default 0.5B student with Qwen2.5-7B teacher (+2.21 on Banking77, +1.95 on HWU64); 1.5B student (+2.82 on Banking77, +3.44 on HWU64); stronger teacher on Banking77 with DianJin-32B (+1.03, 80.30 to 81.33); stronger teacher on HWU64 with Qwen3-32B (+2.32, 84.39 to 86.71).
- Routing rule ablation (Table 3): Preferentially routing low-perplexity samples to GRPO early gives 80.26 on Banking77 and 85.04 on HWU64, versus 78.31 and 84.11 for random routing and 77.27 and 83.46 for high-perplexity-first routing.
- Schedule ablation (Table 4, Banking77): Cosine progressive routing gives 80.26, linear gives 79.58, quadratic gives 79.19, and globally synchronized OPD-to-GRPO gives 78.05.
- Optimization dynamics (Figure 2): PIVOT is reported to achieve higher final test accuracy with substantially smoother reward, entropy, and completion-length dynamics, whereas the globally synchronized baseline shows unstable behavior near its transition stage, including abrupt completion-length increase and entropy spikes.
- Perplexity bucket analysis (Figure 3): High-perplexity samples show substantially larger relative pass@32 improvement during OPD training, while low-perplexity samples saturate much earlier.
- Data augmentation comparison (Table 7, Banking77): From 5,650 candidate augmented examples, 5,119 were retained after teacher self-filtering; fine-tuning on them gives 71.75, above SFT-ZeroShot at 60.30 but below continued OPD (72.47), OPD-to-GRPO (78.05), and PIVOT (80.26).
- Compute caveat: The authors state comparisons are step-matched rather than strictly compute-matched, because PIVOT continues teacher-side perplexity evaluation for samples routed to GRPO and therefore incurs additional teacher inference overhead after the baseline transitions.
Methodology in Plain English
The pipeline has two complementary training signals. On-policy distillation (OPD) makes the student generate its own responses and then matches the teacher's token-level distribution along those trajectories via KL divergence, which is meant to build domain knowledge. GRPO then samples a group of responses per input, assigns scalar rewards, normalizes them into group-relative advantages, and applies a clipped policy-ratio objective, which is meant to refine decisions once knowledge is in place.
The new piece is deciding when each individual example should stop being distilled and start being refined. For each input, the student generates a group of K responses on-policy. The frozen teacher scores each response with length-normalized sequence-level perplexity, these are averaged into a prompt-level score, and the scores are then z-score normalized within the current mini-batch so they are comparable across prompts. A cosine schedule grows a transition ratio r(t) from 0 toward 1 over training; at each step, the lowest-scoring fraction of the batch is routed to GRPO and the rest stays on OPD. The batch loss is simply the average of the GRPO loss for routed samples and the OPD loss for the others. The intuition is that low teacher perplexity means the student's rollouts already sit where the teacher can reliably support them, so rewards can take over, whereas high-perplexity samples still need teacher guidance.
Why This Matters
The paper challenges a common assumption in KD-then-RL pipelines: that more imitation monotonically improves the starting point for later preference or reward optimization. The authors argue that prolonged imitation narrows the student's rollout distribution around teacher-preferred responses, which may reduce the diversity of candidate responses that GRPO needs. This reframes the transition from distillation to reinforcement learning as a scheduling question that should be answered per sample rather than globally.
Real-world applications:
- Customer-service intent routing in banking, where labels are scarce and fine-grained intents overlap heavily.
- Multi-domain virtual assistants, where a small on-device or low-cost model must handle many intents.
- Any regulated vertical where a large expert model can be used offline as a teacher but a small model must serve inference.
- Low-label adaptation projects in other professional domains with structured outputs and checkable rewards.
Industry relevance: the method targets the practical setting where teams have an expert model available but limited labeled data and want a compact deployable model. The paper's own limitation section notes the extra teacher inference cost, which matters for teams weighing compute budgets.
Future Directions
- Developing a theoretical account of why teacher-side perplexity reflects refinement readiness, which the authors explicitly say they do not provide.
- Testing whether the same transition dynamics hold for more complex generation tasks, beyond the structured outputs and reward signals used here.
- Running strictly compute-matched evaluations rather than step-matched ones, given the additional teacher inference overhead PIVOT introduces.
- Investigating whether similar distillation-to-refinement interactions appear in other settings that iterate between imitation and downstream preference optimization.
Target Audience
Researchers and practitioners working on knowledge distillation, reinforcement learning from preference or reward signals, and parameter-efficient adaptation of small language models to specialized domains. It is also relevant to applied teams building few-shot intent classification systems who already have an expert teacher model and want a principled way to coordinate distillation with reward-driven fine-tuning. Readers without background in GRPO or on-policy distillation will need to consult the cited prior work first.
Authors’ abstract
Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD$\rightarrow$GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.