Research
Dynamic Temperature Scheduler for Knowledge Distillation
Overview Research area: Knowledge Distillation (KD) in machine learning, specifically the optimization of the temperature hyperparameter used to soften teacher and student output probabilities. Techni

- arXiv
- 2511.13767
- Published
- 2025-11-14
- Authors
- Sibgat Ul Islam, Jawad Ibn Ahad, Fuad Rahman, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman
AI summary
Overview
- Research area: Knowledge Distillation (KD) in machine learning, specifically the optimization of the temperature hyperparameter used to soften teacher and student output probabilities.
- Technical level: Advanced. The paper includes full derivations of the distillation gradient with respect to temperature and an algorithm listing with a momentum-based update rule, though its main idea can be understood without following the calculus.
- Scope: The paper proposes and evaluates Dynamic Temperature Scheduler (DTS), a temperature schedule that adapts to the difference between teacher and student cross-entropy losses, validated across vision datasets (CIFAR-100, Tiny-ImageNet) and NLP tasks (GLUE, Dolly, SelfInst, S-NI, UnNI, Vicuna).
What This Paper Is About
Knowledge Distillation trains a small student model to imitate a large pre-trained teacher by matching softened output probabilities, where a hyperparameter called temperature controls how soft those probabilities are. Almost all prior work fixes this temperature for the entire training run, even though a student's needs change as it learns. The paper's goal is to show that a fixed temperature is suboptimal and to replace it with a scheduler that adjusts temperature automatically using the gap between the teacher's and the student's cross-entropy losses against the true labels.
Key Contributions
- The paper demonstrates that a static temperature is not always ideal for the KD process, since the optimal degree of softening changes as training proceeds.
- It shows that different temperatures are needed in different training phases: softer probabilities early, when the student's predictions are uncertain, and sharper probabilities later, when the student's logits become more discriminative.
- It introduces Dynamic Temperature Scheduler (DTS), which the authors state is the first temperature scheduling method that adapts based on the divergence between teacher and student distributions, combining cosine scheduling, loss-divergence-based scaling, and a smooth momentum update.
- DTS is designed to integrate seamlessly into existing KD frameworks, and is tested on top of KD, DKD, and MLKD in vision, and KD and SeqKD in language generation.
Main Findings
- DTS beats other temperature-scheduling methods on CIFAR-100 (Table II). Against AKD and CTKD with vanilla KD across six teacher-student pairs, KD + DTS scores highest in every pair: VGG13 to VGG8 gives AKD 67.81, CTKD 72.02, KD + DTS 72.77; ResNet56 to ResNet20 gives 68.13, 69.51, and 70.98; ResNet110 to ResNet32 gives 70.34, 72.05, and 72.20.
- Gains over the KD baseline on same-architecture pairs (Table III). Adding DTS to vanilla KD improves Top-1 accuracy by +0.11, +0.75, +0.30, +0.51, +2.38, and +0.45 across the six teacher-student configurations, with the largest single gain on the VGG13 to VGG8 pair. Adding DTS to DKD gives gains of +0.01, +0.43, +0.12, +0.07, +0.06, and +0.38, and adding it to MLKD gives +0.22, +0.13, +0.26, +0.40, +0.55, and +0.22.
- Stronger gains on cross-architecture pairs (Table IV). KD + DTS improves on KD by +0.28, +0.56, +1.50, +1.42, and +1.61. The authors note that DTS depends heavily on the scheduling range, and that cross-architecture distillation benefits from a lower range.
- Improvements carry over to Tiny-ImageNet (Table V). With CNN-based pairs, ResNet34 to ResNet18 rises from 57.82 (KD) to 60.75 (KD + DTS) at a range of 4 to 2, and ResNet50 to MN-V1 rises from 56.84 to 61.18 at 4 to 2. With ViT-based students and ResNet-50 as teacher, DeiT-Ti rises from 68.42 to 68.58 (4 to 2), T2T-ViT-7 from 64.03 to 64.60 (4 to 2), PvT-Ti from 72.01 to 72.69 (3 to 1), and PiT-Ti from 74.04 to 74.18 (3 to 1).
- Generative language model results are mixed in the reported numbers (Table VI). With a GPT-2 1.5B teacher and a 120M student measured by ROUGE-L, KD + DTS improves on KD on Self (9.82 vs 9.74), S-NI (16.16 vs 14.88), UnNI (19.66 vs 18.46), and Vicuna (13.95 vs 13.92), but reports 21.48 on Dolly versus 21.88 for KD. SeqKD + DTS improves on SeqKD on Self (10.89 vs 9.47) and UnNI (18.77 vs 17.37) but is lower on Dolly (21.30 vs 22.27), S-NI (15.54 vs 15.52 is essentially flat but marginally higher), and Vicuna (14.15 vs 14.90). With an OPT 1.3B teacher and 125M student, KD + DTS improves on KD on all five columns (19.91 vs 19.41, 8.46 vs 8.06, 15.64 vs 15.30, 19.66 vs 17.52, 14.05 vs 13.65), and SeqKD + DTS improves on SeqKD on Dolly (20.33 vs 20.14), Self (8.95 vs 8.24), and S-NI (16.74 vs 15.52), while being lower on UnNI (18.25 vs 17.98 is higher; Vicuna 13.58 vs 14.64 is lower). The paper states that DTS improves existing NLP distillation methods in all cases; the table values as printed do not show improvement in every individual cell.
- The gradient argument for scheduling. Because the gradient of the KD loss with respect to student logits scales as 1/T times the difference between student and teacher probabilities, a very low temperature makes gradients large and destabilizes early training, while a very high temperature makes both distributions approach uniform and the gradients vanish late in training.
- Temperature range is a real design decision (ablation). The authors report that a range of T_max = 8 to T_min = 4 works well for same-architecture distillation (Table III) and T_max = 3 to T_min = 1 helps cross-architecture distillation (Table IV). Figure 2 shows ResNet20 distilled by ResNet56 and ResNet110 over 50-epoch CIFAR-100 training across ranges 3 to 1, 4 to 2, 6 to 4, 8 to 4, and 11 to 9.
- Clamping is necessary. The paper reports that without the clamping step, the temperature changes abruptly and the student performs worse.
Methodology in Plain English
The authors start from the observation that a temperature value acts as a volume knob on the learning signal. Set it too low early on, when the student's logits are near zero and its probabilities are nearly uniform, and the 1/T factor inflates small teacher-student disagreements into oversized updates. Set it too high late in training, when the student already has confident logits, and the teacher and student distributions both flatten toward uniform, shrinking the gradient to nearly nothing.
DTS handles both cases automatically. At each epoch it computes a training progress value p, the ratio of the current epoch to the total number of epochs. It maps p through a cosine curve, S(p) = λ(1 + cos(π·p)) with λ set to 0.5, which starts at 1 and decays to 0. It computes the cross-entropy loss of both the teacher and the student against the true labels, takes their difference, and converts that difference into a scaling coefficient α. When α exceeds 1, it multiplies the cosine value by α to push the temperature up; otherwise the cosine term alone is used. The tentative temperature is clipped between a minimum and a maximum bound, and then the actual temperature for the epoch is updated with a momentum rule, a weighted blend of the previous temperature and the new target with a default momentum coefficient of 0.9. That final smoothing is what keeps the schedule from jumping around and destabilizing training.
The paper's evaluation is broad by design: the same scheduler is dropped into vanilla KD, DKD, and MLKD on CIFAR-100 and Tiny-ImageNet using ResNet, VGG, MobileNet, ShuffleNet, and several ViT variants including DeiT, PiT, PvT, and T2T-ViT; and into KD and SeqKD for GPT-2 and OPT generation, plus TinyBERT distilled from BERT_base_uncased on three GLUE tasks (CoLA, MRPC, RTE). Vision training used SGD, 100 epochs on CIFAR-100 with a 0.01 initial learning rate for MobileNets and ShuffleNets and 0.1 for other models, decaying by 0.1 at the 63rd, 87th, and 92nd epochs; Tiny-ImageNet used 50 epochs with an initial learning rate of 0.2 decayed by 10x at epochs 15, 30, and 45, and results are averaged over 3 trials. NLP experiments used T_init = T_max = 4 and T_min = 2.
Why This Matters
- Impact on research. The paper challenges a default that has gone largely unexamined since Hinton et al. introduced KD: that temperature is a single fixed constant. It also argues directly against prior schedulers such as AKD (two-stage, not dynamic), CTKD (uses learnable temperature modules and adversarial learning), and RLKD (a reinforcement learning pipeline with computational overhead), presenting DTS as requiring no learnable module and no multi-stage training.
- Lowering the tuning cost of KD pipelines. Temperature is normally chosen by extensive grid search that varies by dataset and by model architecture. An automatic schedule reduces that manual search and makes distillation more portable across teacher-student pairs.
- Real-world applications:
- Deploying smaller vision models on phones, cameras, and edge devices by distilling a large CNN or ViT teacher into a compact student without hand-tuning per deployment.
- Compressing generative language models such as GPT-2 1.5B or OPT 1.3B down to roughly 120M-125M parameter students for cheaper inference, since those exact pairings are evaluated here.
- Producing compact task models for GLUE-style classification tasks, demonstrated with TinyBERT-4L-312D distilled from BERT_base_uncased.
- Improving any existing in-house distillation pipeline, since DTS is described as a drop-in addition to frameworks such as KD, DKD, MLKD, and SeqKD rather than a replacement for them.
- Industry relevance. Model compression is a direct cost lever for inference-serving budgets, and a scheduler that needs no extra learned modules or training stages is easier to insert into production training code than an adversarial or reinforcement-learning alternative.
Future Directions
- Assigning different temperatures to the student and the teacher separately, which the authors suggest could help bridge the performance gap caused by architectural differences.
- Instance-wise temperature adjustment, adapting temperature to individual sample difficulty so that easier samples use lower temperatures and challenging samples use higher ones, which would also address the authors' concern that batch-wise adjustment can be destabilized by hard batches.
- Reducing the tuning burden of the temperature range. The authors list the fact that choosing T_max to T_min requires tuning, with different ranges working better for same-architecture versus cross-architecture pairs, as an explicit limitation of the current method.
- Closing the gap between the reported DTS results and the strongest feature-distillation baselines, such as SimKD's 78.08 on the ResNet32x4 to ResNet8x4 pair, and clarifying the mixed generative-model results where the paper's claim of improvement in all NLP cases is not visible in every cell of Table VI.
Target Audience
Machine learning researchers and practitioners working on model compression and efficient inference, particularly those already familiar with Knowledge Distillation and its temperature hyperparameter. It is most useful to engineers who need to distill models across mismatched architectures and want to avoid grid-searching temperature per teacher-student pair, and to researchers interested in training-dynamics-based hyperparameter scheduling. Readers without background in KD will find the gradient derivations dense, but the core scheduling idea is described in enough plain detail to follow.
Authors’ abstract
Knowledge Distillation (KD) trains a smaller student model using a large, pre-trained teacher model, with temperature as a key hyperparameter controlling the softness of output probabilities. Traditional methods use a fixed temperature throughout training, which is suboptimal. Moreover, architectural differences between teacher and student often result in mismatched logit magnitudes. We demonstrate that students benefit from softer probabilities early in training but require sharper probabilities in later stages. We introduce Dynamic Temperature Scheduler (DTS), which adjusts temperature dynamically based on the cross-entropy loss gap between teacher and student. To our knowledge, this is the first temperature scheduling method that adapts based on the divergence between teacher and student distributions. Our method integrates seamlessly with existing KD frameworks. We validate DTS across multiple KD strategies on vision (CIFAR-100, Tiny-ImageNet) and NLP tasks (GLUE, Dolly, SelfIns, UnNI, S-NI), consistently outperforming static-temperature baselines. Code is available at https://github.com/Sibgat-Ul/DTS.