Research
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
Overview Research area: Natural Language Processing — knowledge distillation of large reasoning models (Chain-of-Thought distillation) and reward modeling for reinforcement learning, with an industria

- arXiv
- 2511.01354
- Published
- 2025-11-03
- Authors
- Wenrui Cai, Chengyu Wang, Junbing Yan, Jun Huang, Xiangzhong Fang
AI summary
Overview
Research area: Natural Language Processing — knowledge distillation of large reasoning models (Chain-of-Thought distillation) and reward modeling for reinforcement learning, with an industrial deployment angle.
Technical level: Intermediate. The paper assumes familiarity with supervised fine-tuning, Chain-of-Thought reasoning, teacher-student distillation, and GRPO-style reinforcement learning, but the pipeline description itself is high-level and process-oriented.
Scope: The paper extends the DistilQwen family with four distilled model series — one slow-thinking series, two adaptive-thinking series, and a series of distilled reward models — and evaluates them on mathematics, coding, and general question-answering benchmarks, plus integration with the Alibaba Cloud PAI platform.
What This Paper Is About
Industries need reasoning models that are both accurate and fast, but the strongest large reasoning models are too expensive to serve. The authors attack this gap by distilling Chain-of-Thought reasoning from very large teachers (DeepSeek-R1, DeepSeek-R1-0528, and QwQ-32B) into small Qwen2.5 and Qwen3 students, and by additionally distilling reward signals that can drive later reinforcement learning. The goal is a practical, released family of small reasoning models that trade off accuracy and inference cost in a controllable way.
Key Contributions
-
The DistilQwen2.5-R1 slow-thinking series in 3B, 7B, 14B, and 32B parameter scales, trained on CoTs that are generated, difficulty-scored, rewritten, and verified by DeepSeek-R1, with SFT curriculum learning followed by DPO.
-
Two adaptive-thinking series. DistilQwen-ThoughtX (7B and 32B, Qwen2.5 students, trained with the dataset from Cai et al. 2025a) and DistilQwen-ThoughtY (4B, 8B, and 32B, Qwen3 students with thinking modes enabled, trained on the previous series' dataset plus a subset of 365K CoTs generated from DeepSeek-R1-0528). These adjust CoT length to problem difficulty rather than always reasoning at maximum length.
-
Distilled reward models. Two reward predictors, each initialized from Qwen2.5-7B-Instruct, that estimate a Reasoning Verbosity (RV) score and a Cognitive Difficulty (CD) score, allowing GRPO reinforcement learning to be guided by teacher-derived knowledge without running the original teachers as reward models.
-
An industrial pipeline and release. A Data Source Collector, elastic teacher inference, LLM-based CoT processors, and a training stack, with resources released in the EasyDistill toolkit and models integrated into the Alibaba Cloud PAI platform.
Main Findings
-
Slow-thinking models beat their bases and same-teacher baselines. DistilQwen2.5-7B-R1 (105K training set) scores 43.3 on AIME2024, 88.4 on MATH-500, 42.9 on GPQA Diamond, and 46.4 on LiveCodeBench V2, versus 10.0 / 73.6 / 33.3 / 30.7 for Qwen2.5-7B-Instruct and 31.3 / 83.0 / 42.4 / 39.9 for OpenThinker-7B, which uses the same source reasoning problems and the same teacher.
-
Gains hold across scales. DistilQwen2.5-3B-R1 reaches 16.7 / 70.0 / 34.3 / 18.0 (base: 6.7 / 62.6 / 32.8 / 11.3); DistilQwen2.5-14B-R1 reaches 46.7 / 90.8 / 51.5 / 54.4 (base: 16.7 / 78.2 / 43.4 / 37.4); DistilQwen2.5-32B-R1 reaches 70.0 / 93.8 / 62.1 / 66.0 (base: 16.7 / 81.4 / 45.5 / 47.3; OpenThinker-32B: 66.0 / 90.6 / 61.6 / 68.9).
-
Adaptive-thinking models surpass slow-thinking ones. DistilQwen-ThoughtX-7B scores 56.7 / 90.2 / 50.0 / 56.8 against Qwen2.5-7B-Instruct at 10.0 / 73.6 / 33.3 / 30.7. DistilQwen-ThoughtX-32B scores 80.0 / 92.6 / 64.0 / 73.4 against Qwen2.5-32B-Instruct at 16.67 / 81.4 / 45.5 / 47.3.
-
ThoughtY improves on ThoughtX. DistilQwen-ThoughtY-32B scores 90.0 on AIME2024 versus 76.7 for Qwen3-32B in thinking mode, and 76.3 on LiveCodeBench V2 versus 72.2. ThoughtY-8B scores 78.1 on LiveCodeBench V2 versus 62.8 for Qwen3-8B thinking mode.
-
Adaptive models produce difficulty-appropriate CoT lengths. On GSM8K the models produce shorter CoTs than the slow-thinking counterparts (ThoughtX-7B: 834.36; ThoughtY-8B: 844.23; 7B-R1: 1223.61), while on AIME2024 they produce longer ones (ThoughtX-7B: 14597.96; ThoughtY-8B: 15632.85; 7B-R1: 12856.23). The same pattern holds at 32B (ThoughtX-32B: 742.04 / 5927.32 / 16387.53; ThoughtY-32B: 723.18 / 5723.08 / 17231.84; 32B-R1: 1178.92 / 6434.50 / 13583.19).
-
Distilled reward models improve RL. With Qwen2.5-7B-Instruct and 10K sampled mathematical problems for RL and no CoT-based SFT, MATH500 / AIME2024 results are 73.6 / 10.0 raw, 78.8 / 13.3 for vanilla GRPO, 79.0 / 13.3 for GRPO+RV, 80.8 / 16.7 for GRPO+CD, and 81.4 / 20.0 for GRPO+RV+CD.
-
Data volume explains the adaptive-thinking advantage. Adaptive-thinking models sample from over 2 million CoTs, giving training sets of at least 500K, while slow-thinking models train on about 100K data points. The authors state the slow-thinking recipe remains preferable when training data is inherently limited, because adaptive sampling would further shrink the usable set.
-
Inference-time scaling helps. Pass@K results show that increasing the number of reasoning attempts K produces significant accuracy gains for the DistilQwen2.5-R1 models, with the 7B model trending steeply upward on MATH500 and GPQA Diamond and approaching the 32B model.
Methodology in Plain English
The team first collects Chain-of-Thought datasets from sources such as Hugging Face and ModelScope, covering mathematics, code, and science — including resources named as OpenThoughts2, DeepMath-103K, and OpenCodeReasoning — and re-samples them so task types are balanced. Teacher models are not called through third-party APIs; instead, DeepSeek-R1, DeepSeek-R1-0528, and QwQ-32B are served on the authors' own clusters, with each server holding eight NVIDIA H20 GPUs of 96GB each, and the number of inference nodes per model can be scaled elastically with demand.
Two CoT processors then clean and reshape that data. The slow-thinking processor uses DeepSeek-R1 to generate multiple CoTs per problem at varying temperatures, scores each CoT's difficulty as easy, medium, or hard, and rewrites and verifies only the easy and hard CoTs in a single pass, discarding incorrect ones. The idea is that students learn best from medium-level CoTs, and that at least one suitable CoT usually survives because many are generated per problem. The adaptive-thinking processor adds two more scorers: Reasoning Verbosity, which checks whether a CoT is appropriately long for the problem, and Cognitive Difficulty, which checks whether it matches the student's capacity. A target-aware sampler then picks an optimal subset of CoTs for the given student model.
Training uses supervised fine-tuning with curriculum learning: medium-level CoTs first for stable convergence, then progressively harder samples. Slow-thinking models are trained for three epochs on medium-level CoTs, then on harder examples, then refined with DPO. Separately, the RV and CD scores are used to train two lightweight reward predictors initialized from Qwen2.5-7B-Instruct, which stand in for the large teachers during GRPO reinforcement learning. The overall GRPO reward combines format and accuracy rewards with RV and CD penalty terms that are zero when a predicted score lies inside a designated interval and increase linearly with distance outside it, weighted by tunable coefficients.
Why This Matters
Impact on research. The paper shows that careful data curation — difficulty scoring, verification, and verbosity/cognitive-difficulty control — matters as much as raw distillation scale, and it opens a route to reinforcement learning that does not require the original ultra-large teachers as reward models. It also documents a negative-but-useful result: when data is scarce, the more selective adaptive recipe can hurt by shrinking the training pool.
Real-world applications:
- Mathematical and scientific problem solving in assistants that must run at low cost.
- Code generation and algorithmic assistance, evaluated via LiveCodeBench V2.
- Reinforcement-learning pipelines that need cheap, local reward signals instead of expensive teacher calls.
- Multi-agent and decision-making systems where small reasoning models are deployed at scale.
Industry relevance. The models are deployed on the Alibaba Cloud PAI platform covering training, evaluation, compression, and deployment, and are exposed through RESTful APIs compatible with the OpenAI format. All reasoning models are also released open-source, with resources in the EasyDistill toolkit.
Future Directions
-
Extending RL-enhanced lightweight reasoning models, which the authors explicitly flag as future work beyond the knowledge distillation techniques presented here.
-
Continuing to refresh the data pipeline as stronger public teachers appear; the jump from ThoughtX to ThoughtY demonstrates the payoff of collecting higher-quality CoTs, and the ThoughtY subset of 365K CoTs from DeepSeek-R1-0528 is described as to be released.
-
Addressing limitations in real-world contexts, since benchmark performance may not capture highly specific or dynamic operational complexity.
-
Mitigating bias and error propagation, because the reward models depend on distilled knowledge and may inherit biases or errors from the teacher models; the authors also call out data privacy, security, and regulatory compliance as implementation concerns.
Target Audience
Practitioners and researchers working on knowledge distillation, small reasoning models, or RL fine-tuning will get the most value, especially those who need concrete benchmark numbers and a described data-curation pipeline rather than purely theoretical results. Engineers evaluating deployment on cloud AI platforms, and teams deciding between a simpler high-accuracy distillation recipe and a more data-hungry adaptive one, will find the data-volume discussion directly actionable. Beginners can follow the high-level pipeline but may need background in Chain-of-Thought reasoning and GRPO to interpret the reward formulation.
Authors’ abstract
Recently, the demand for small and efficient reasoning models to support real-world applications has driven the development of knowledge distillation techniques that balance reasoning performance and inference speed. In this paper, we further extend the DistilQwen model family, initialized from the Qwen models, by introducing four model series specifically designed to meet industrial requirements. The distilled model collection comprises: (1) slow-thinking models, optimized for reasoning tasks that require high accuracy; (2) two series of adaptive-thinking models, which dynamically adjust reasoning strategies based on input tasks to maximize efficiency across diverse scenarios; and (3) distilled reward models, which enable further reinforcement learning of reasoning models using distilled knowledge. Comprehensive evaluations across multiple benchmarks demonstrate both high inference efficiency and strong reasoning performance for these models, as well as the practical utility of distilled reward models. We further show that these models support industry practitioners by providing scalable training and inference functionalities on the Alibaba Cloud PAI (Platform for Artificial Intelligence) platform.