Skip to content
AI.info

Research

GIFT: Reconciling Post-Training Objectives via Variational Finite-Temperature Gibbs Initialization

The prevailing post-training paradigm for Large Reasoning Models (LRMs)---Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL)-suffers from an intrinsic optimization mismatch: the rigi

arXiv
2601.09233
Published
2026-01-14
Authors
Zhengyang Zhao, Lu Ma, Yizhen Jiang, Xiaochen Ma, Zimo Meng, Chengyu Shen, Lexiang Tang, Haoze Sun, Peng Pei, Wentao Zhang

Authors’ abstract

The prevailing post-training paradigm for Large Reasoning Models (LRMs)---Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL)-suffers from an intrinsic optimization mismatch: the rigid likelihood maximization in SFT induces distributional collapse, thereby exhausting the exploration space necessary for subsequent RL. Motivated by the Gibbs optimum of KL-regularized RL, we derive a token-level variational surrogate that makes the SFT target structurally compatible with the subsequent RL stage, and propose Gibbs Initialization with Finite Temperature (GIFT). Standard SFT emerges as a degenerate zero-temperature limit of this surrogate, while a finite temperature preserves structural diversity. Our experiments demonstrate that GIFT outperforms standard SFT and surpasses competitive baselines when utilized for RL initialization. Our code is available at https://github.com/zzy1127/GIFT.

Read the original paper