Research
Smaller Models, Better Rejects: Preference Distillation Scaling
Overview Research area: Preference-based knowledge distillation and post-training data design for large language models, specifically how the rejected response in a DPO preference pair should be const

- arXiv
- 2609.38987
- Published
- 2026-09-30
- Authors
- Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov, Kayhan Behdin, Zhengze Zhou, Ran He, Alborz Geramifard, Rohit Jain, Zhe Zhao
AI summary
Overview
- Research area: Preference-based knowledge distillation and post-training data design for large language models, specifically how the rejected response in a DPO preference pair should be constructed.
- Technical level: Advanced. The paper combines a large empirical scaling study (7B–72B students, code and math) with a theoretical analysis using a linearized feature model of Direct Preference Optimization and a finite-horizon utility bound.
- Scope: The paper shows that frozen models smaller than the student generate better rejected responses than the student's own outputs, and it explains this through two measurable properties of useful rejects: task structure and limited coupling to the DPO reference policy.
What This Paper Is About
Preference distillation usually builds training pairs by treating a teacher response as the preferred answer and the student's own response as the rejected answer, on the assumption that a model's own failures are its most informative negatives and that reject generation must scale along with the student. The authors test both assumptions and find neither holds: frozen "smaller Base" models with strictly fewer parameters than the student produce rejects that train stronger students at lower inference cost. The goal is to characterize what property of a reject distribution actually makes it useful, and to derive a practical design principle for constructing rejects.
Key Contributions
- A scaling phenomenon in preference distillation. Across Qwen2.5 students from 7B to 72B and across code generation and mathematical reasoning, rejects from smaller frozen Base models require less inference compute yet train stronger students than either vanilla Self or post-SeqKD Self rejects.
- A theoretical framing of reject construction as inverse data design. The authors use a linearized feature model of DPO and derive a finite-horizon transfer bound that characterizes a favorable region of reject distributions via two coordinates: transfer mass (a_Q) and adverse share (c_Q).
- Three construction interventions derived from that bound and tested empirically: randomized source mixtures (H1), prompt reassignment and lexical permutation (H2), and reference-likelihood reselection (H3).
- Two observable properties of useful rejects. Task-relevant structure and limited coupling to the reference policy, yielding the design principle that effective rejects preserve task structure while limiting coupling to the reference policy, and smaller frozen models supply both at low cost.
Main Findings
- Smaller Base rejects beat Self rejects at every tested scale. All 18 smaller Base configurations in the code scaling matrix outperform both post-SeqKD Self and vanilla Self on avg@4. For example, rejects generated by Qwen2.5-3B-Instruct improve the 14B student's avg@4 over vanilla Self by 3.1 points on code generation and 1.9 points on mathematical reasoning.
- The advantage is cheaper. On the code preference prompts, estimated generation compute for smaller Base sources is 0.9% to 50.1% of vanilla Self at the same student scale, corresponding to a 2.0 times to 108.3 times reduction.
- Mixing rejects moves performance monotonically with the smaller model's share. On a matched 14B training population, code avg@4 rises from 59.73% to 63.28% as the share of Llama-1B Base rejects increases from 0% to 100%. The Qwen2.5-3B mixture follows the same ordering from 59.58% to 62.63%, and a 72B midpoint lies between its two pure source endpoints.
- Gibberish rejects add little. Length-matched artificial rejects closely track Continued-SFT (training on the chosen responses alone) at 14B, 32B, and 72B, with a larger gain at 7B.
- Prompt correspondence is not required. Prompt-reassigned Base rejects — real reject text attached to different prompts — improve over Gibberish at every student scale, with the largest gain for the 7B student at 1.28 points of avg@4.
- Code-domain lexical structure alone carries corrective value. Permuting tokens within prompt-reassigned rejects destroys token order, Python syntax, and prompt correspondence, yet Lexical still outperforms Gibberish at every scale. The largest gain is for the 14B student, at 0.73 points of avg@4.
- Reference likelihood organizes the natural source boundary. Scoring each natural reject source under the student's exact SeqKD reference, higher reference likelihood is associated with lower downstream avg@4 at every student scale, with the strongest relationship for the 7B student (Pearson correlation of −0.97). Smaller Base sources combine lower coupling with higher utility; vanilla Self and post-SeqKD Self combine higher coupling with lower utility.
- Reselection by likelihood is causal, not just a proxy. From a fixed bank of eight candidates per prompt, selecting lower-likelihood rejects (Repelled) produces the strongest endpoint for every source, while selecting higher-likelihood rejects (Attracted) reduces utility for each smaller Base source. The largest separation is for the 1.5B source, where Repelled reaches 63.14% avg@4 versus 61.88% for Attracted, a gain of 1.26 points.
- Optimization diagnostics do not explain the effect. Reject fields, reward margins, likelihood trajectories, and RC-DPO do not recover the smaller Base over Self ordering (Appendix E).
Methodology in Plain English
The setup is the two-stage SODA pipeline. A student first learns from execution-verified teacher responses through sequence-level knowledge distillation (SeqKD), and then undergoes DPO on a separate preference set. Everything is held fixed within a student scale — prompts, chosen responses, the SeqKD initialization used as the DPO reference, the objective, optimizer, training schedule, budget, and evaluation protocol. Only the reject construction changes.
The comparison is between three constructions: "vanilla Self" (the matching instruction-tuned model at the student's scale), "post-SeqKD Self" (the reference policy that initializes DPO), and "smaller Base" (a frozen vanilla model with strictly fewer parameters than the student).
For the theory, the authors linearize the network around the shared DPO initialization and treat a reject as a set of features relative to the chosen response. Rejects are scored by how far they move the initial update along the direction that improves task utility, decomposed into a transfer mass and an adverse share — the fraction of that movement pointing the wrong way. A finite-horizon bound states that utility gain is at least a coefficient times net transfer minus an error term that scales with the square of total step size, which defines a favorable region of reject distributions.
Three interventions follow. First, randomly mix rejects from a smaller model and from the student-scale model, since net transfer should be affine in the mixing weights. Second, reassign rejects to other prompts and permute their code tokens, since a prompt-independent component of the features should survive. Third, reselect candidates from a fixed per-prompt bank by their length-normalized likelihood under the SeqKD reference, repelling from the reference (γ > 0), attracting to it (γ < 0), or sampling natively (γ = 0).
The code line uses KoDCode: SeqKD on 75,752 prompts with execution-verified GPT-4o responses, DPO on a disjoint 24,248-prompt preference split, and evaluation on 5,000 held-out KoDCode problems plus BigCodeBench-Complete and BigCodeBench-Instruct (1,140 problems), HumanEval (164 problems), and MBPP (257 problems). Every reject is verified incorrect by failing at least one unit test. The math line uses verifier-correct DeepSeek-R1 trajectories from OpenR1-Math-220k, with 33,524 prompts for SeqKD and 12,288 preferences for DPO, evaluated on MATH-500 and the 2024 and 2025 AIME problems. The primary in-domain endpoint is avg@4, the mean execution success over four sampled responses. Training uses full-parameter training with AdamW, a cosine schedule with 10% linear warmup, bf16 precision, gradient checkpointing, and DeepSpeed ZeRO-3, with learning rates of 5×10⁻⁶ (SeqKD) and 5×10⁻⁷ (DPO), β = 0.1, 1 epoch, effective batch 256, max sequence length 2,560, and gradient clipping 0.2.
Why This Matters
Impact on research. The paper reframes reject construction from a byproduct of the training pipeline into an independent design variable, and it supplies a theoretical account — a finite-horizon utility bound with two measurable coordinates — for why data that looks "worse" (errors from a weaker model) can produce a better student. It also shows that several natural optimization-level diagnostics fail to explain the ordering, which constrains how future work should attribute gains from preference data.
Real-world applications:
- Cost-efficient post-training pipelines. Reject generation can be delegated to models whose estimated generation compute is 0.9% to 50.1% of a same-scale Self generator, freeing the largest model for chosen-response generation only.
- Distillation where only black-box outputs are available. Because prompts remain a fixed resource, the approach changes only the reject source, making it compatible with fixed prompt sets and fixed chosen-response pools.
- Data-quality auditing. Reference likelihood under the student's own initialization is a cheap, observable signal that separates stronger from weaker reject sources and can be used to filter or reselect existing candidate banks.
- Multi-task post-training. The two identified properties — task structure and low reference coupling — give concrete guidance for constructing rejects in code, math, and potentially other verifiable domains.
Industry relevance. The result directly reduces the cost of the reject half of preference data, which is often the more expensive half when sampling from large models. The likelihood-based reselection method requires only scoring existing candidates under the frozen reference, so it can be layered onto existing DPO pipelines without changing the generator, the student initialization, or the optimization protocol.
Future Directions
- Does the smaller-is-better pattern extend beyond 72B and beyond Qwen2.5 and Llama families? The study covers students from 7B to 72B and notes that the best source varies with student and task; the behavior at larger scales and other model families is not established.
- How far can the small-source advantage go? The paper tests strictly smaller frozen Base sources with fewer parameters than the student, but does not report a lower bound on how small a reject generator can be before utility degrades.
- What exactly is the task-structure signal? The results show that prompt-reassigned and lexically permuted rejects beat gibberish, and an AST-based intervention is reported in Section D.2 of the paper, but the precise feature-level account of which code-domain structure carries corrective value remains an open question.
- Can reference coupling be exploited more aggressively than reselection? Reselection operates over a fixed bank of eight candidates per prompt; whether generating candidates specifically to be low-likelihood under the reference, or combining repulsion with mixture and reassignment interventions, yields further gains is not reported.
Target Audience
This paper is most useful to machine learning researchers and engineers working on LLM post-training, preference optimization, and knowledge distillation, particularly those building DPO or SODA-style pipelines who control how preference pairs are constructed. It also suits data-centric ML practitioners interested in cheap reject generation, and readers with a background in optimization who want a theoretical framing of why certain negative examples transfer better than others. The empirical results are legible without the theory, but the favorable-region bound and the transfer coordinates require comfort with DPO objectives and linearized model analysis.
Authors’ abstract
Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.