Research
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
Overview Research area: Language model alignment, specifically direct preference optimization (DPO) and the theory of replacing the KL divergence in the RLHF objective with general $f$-divergences. Te

- arXiv
- 2602.06788
- Published
- 2026-02-06
- Authors
- Idan Pipano, Shoham Sabach, Kavosh Asadi, Mohammad Ghavamzadeh
AI summary
Overview
- Research area: Language model alignment, specifically direct preference optimization (DPO) and the theory of replacing the KL divergence in the RLHF objective with general $f$-divergences.
- Technical level: Advanced (requires familiarity with reinforcement learning from human feedback, $f$-divergences, and convex analysis).
- Scope: This paper characterizes exactly which functions $f$ keep the generalized RLHF problem tractable, identifies a second condition that provably guards against probability displacement, and instantiates both conditions in a new loss called SquaredPO.
What This Paper Is About
DPO aligns language models by directly optimizing the RLHF objective: maximize a Bradley-Terry reward while penalizing deviation from a reference policy, originally with a KL divergence. Prior work showed this stays tractable when the KL term is replaced by an $f$-divergence whose generating function $f$ is convex, differentiable, and has an invertible derivative $f'$ with $0 \notin \text{dom}(f')$. This paper asks two questions: what is the complete set of functions $f$ for which the problem remains directly solvable, and among those, which choices avoid the empirical failure mode where the probabilities of both winner and loser responses collapse toward zero.
Key Contributions
-
Full characterization of tractable functions. The paper defines a function $f$ as DPO-inducing if substituting any optimal solution of the generalized RLHF problem into the Bradley-Terry model yields the $f$-DPO loss. It proves this property holds for a strictly larger class than the convex functions used by prior work. Under the additional assumption that $\lim_{t \to 0^{+}} f'(t)$ exists (possibly $\pm\infty$), Corollary 1 states that $f$ is DPO-inducing if and only if $\lim_{t \to 0^{+}} f'(t) = -\infty$.
-
A provable criterion against probability displacement. The paper introduces the condition $\arg\min_{t \in \mathbb{R}{+}} f(t) \geq 1$ and calls any $f$ satisfying it displacement-resistant. Lemma 2 shows that if $f$ is DPO-inducing with a unique global minimum at $c \in (0,1]$, then under a mild assumption any optimal solution to a restricted variant of the problem satisfies $\pi{\theta}(y \mid x) \leq c \cdot \pi_{\text{ref}}(y \mid x)$ for in-sample responses, i.e. in-sample probabilities shrink by at least a factor of $c$.
-
A new loss, SquaredPO. The paper proposes $f_{\text{SquaredPO}}(t) \coloneqq (\log t)^{2}/2$, deliberately non-convex, which is both DPO-inducing and displacement-resistant. Substituting $f'{\text{SquaredPO}}(t) = (\log t)/t$ gives a loss that can be read as "DPO with adaptive $\beta$-s," where $\beta{\theta}(y,x) \coloneqq \beta / \frac{\pi_{\theta}(y \mid x)}{\pi_{\text{ref}}(y \mid x)}$ rises as a response's probability falls, strengthening regularization on that response.
-
Empirical validation and a new empirical observation. The paper shows SquaredPO is robust to over-optimization, competitive on standard benchmarks, and empirically mitigates likelihood displacement, and it reports a previously unreported "monotonicity" phenomenon in DPO training dynamics.
Main Findings
-
Convexity is not required. Contrary to the standard motivation for convex $f$ (that Jensen's inequality plus $f(1)=0$ guarantees non-negativity), the paper shows the summation in the relevant term runs over a strict subset of the response space in one of the associated problems, so Jensen's inequality does not apply and convexity no longer guarantees non-negativity. The paper states it is the first to explore $f$-DPO with non-convex functions $f$.
-
Failure to satisfy $\lim_{t \to 0^{+}} f'(t) = -\infty$ means intractability. The authors note this has a two-fold significance: it broadens the usable family of $f$, and it shows that violating the condition makes the optimization problem intractable.
-
$f$-DPO also solves a conceptually problematic objective. Lemma 1 shows that for any DPO-inducing $f$, substituting an optimal solution of Problem (7) — where the regularization sums only over in-sample responses $y \in S_x$ rather than the whole response space — into the Bradley-Terry model yields the same $f$-DPO loss. This holds regardless of convexity.
-
The global minimum of $f$ controls displacement. Applying Lemma 2 to vanilla DPO, whose generator is $f_{\text{KL}}(t) = t \log t$ with $\arg\min_{t \in \mathbb{R}{+}} f{\text{KL}}(t) = e^{-1}$ (given as 0.36788), implies in-sample probabilities decrease by at least a factor of $e^{-1}$. The paper states the displacement result for general $f$ generalizes what prior work (Asadi et al., 2025) had shown only for DPO.
-
$\chi$PO sits between DPO and SquaredPO in displacement risk. The $\chi$PO generator $f(t) = \tfrac{1}{2}(t-1)^2 + t\log t$ attains its global minimum at $W(1) \approx 0.56714$ (where $W$ is the Lambert $W$ function). Since $e^{-1} < W(1) \leq 1$, theory predicts $\chi$PO is less displacement-prone than vanilla DPO but more so than SquaredPO, whose minimum is at $c = 1$.
-
SquaredPO is robust to over-optimization on TL;DR. Win rate against the base model Meta-Llama-3-8B-Instruct: at epoch 1, SquaredPO 50.8 ± 0.7, $\chi$PO 51.2 ± 0.8, DPO 51.8 ± 1.0; at epoch 2, 50.6 ± 1.1, 48.9 ± 1.1, 45.0 ± 1.1; at epoch 4, 51.0 ± 0.7, 48.3 ± 1.1, 34.7 ± 1.3. DPO's win rate drops substantially below 50% already in the second epoch, and the gap grows with epochs.
-
DPO edges out SquaredPO on the first epoch and on benchmarks. The head-to-head win rate of SquaredPO against DPO shows statistically significant improvements for SquaredPO in regimes prone to over-optimization (two epochs or more), while DPO outperforms SquaredPO at epoch 1 (not statistically significant). On standard benchmarks after one epoch: AlpacaEval 2 length-controlled win rate 29.2 ± 0.4 (SquaredPO), 29.3 ± 0.6 ($\chi$PO), 29.6 ± 0.4 (DPO); raw win rate 24.5 ± 0.3, 24.3 ± 0.4, 24.8 ± 0.4; MT-Bench score (1–10) 7.924, 7.900, 7.925. The authors note hyperparameters were standard DPO choices (e.g. $\beta = 0.01$) and were not tuned for SquaredPO.
-
Displacement is smaller and less extreme under SquaredPO. Histograms of chosen log-ratios $\log(\pi_{\theta}(y_w \mid x)/\pi_{\text{ref}}(y_w \mid x))$ after one epoch show the magnitude of decrease is smaller under SquaredPO, and the most extreme decreases seen under DPO are absent. The right panel of Figure 3 shows the mean and median chosen log-ratios per epoch, with DPO's decrease described as far more radical. Appendix results over 1, 2, 3 and 4 epochs indicate the mitigation grows more pronounced with additional training.
-
A monotonicity phenomenon in DPO. In this experimental setting, 99.63% of chosen responses whose probability decreased in the first epoch continued decreasing monotonically in the three subsequent epochs; under SquaredPO only 4.21% did. The authors state they are, to their knowledge, the first to report this monotonicity phenomenon checked on a per-winner basis. This is consistent with the mechanism of SquaredPO: when a winner suffers displacement, its effective regularization coefficient $\beta_{\theta}(y_w, x)$ increases, encouraging its probability to return toward $\pi_{\text{ref}}(y_w \mid x)$.
Methodology in Plain English
The authors work analytically first. They define the generalized RLHF problem where the KL penalty is replaced by an $f$-divergence built from a generating function $f$, then ask which choices of $f$ still allow the closed-form optimal policy to be plugged into the Bradley-Terry preference model so that intractable quantities cancel. They name that property DPO-inducing and prove a complete characterization of it. They connect DPO-inducing to a second property they call interior-inducing — whether optimal solutions assign non-zero probability to all responses — and use that equivalence in the proof.
They then separate the problem from the full response space and study a restricted version that only regularizes over responses that actually appear in the preference dataset. By analyzing where $f$ attains its minimum, they derive a bound on how much in-sample probabilities can shrink, which converts into a clean, easy-to-check criterion: the minimum of $f$ should be at or above 1. They audit several classical generators against both criteria in a Venn diagram (convex, DPO-inducing, displacement-resistant) and pick $f(t) = (\log t)^2/2$ as a non-convex function lying in the intersection.
For experiments, they fine-tune Meta-Llama-3-8B-Instruct on the TL;DR preference dataset (the version preprocessed by von Werra et al. (2020)) for four epochs using LoRA. Evaluation proceeds on three axes: chosen log-ratios across checkpoints to measure displacement, GPT-4-judged win rates on a 512-example sample of the validation split (with checkpoints at epochs in {1, 2, 4}), and out-of-distribution benchmarks AlpacaEval 2 (805 single-turn prompts, judged by GPT-4, reporting raw and length-controlled win rates) and MT-Bench (80 multi-turn questions, scored 1–10 by GPT-4, averaged over eight categories). AlpacaEval numbers are averaged over 10 seeds with confidence intervals; Figure 2 reports error bars over 10 seeds. OpenAI's gpt-4o is used instead of the default judge gpt-4-1106-preview for reduced API costs. Baselines are DPO and $\chi$PO.
Why This Matters
Impact on research. The paper turns a loose design space into two checkable mathematical conditions: $\lim_{t \to 0^{+}} f'(t) = -\infty$ for tractability and $\arg\min_{t \in \mathbb{R}_{+}} f(t) \geq 1$ for displacement resistance. It also shows that the usual justification for requiring convex $f$ does not survive in this setting, which opens the $f$-DPO family to non-convex generators. The reported monotonicity phenomenon (99.63% of early-decreasing winners continuing to decrease under DPO) is a new empirical characterization of how displacement evolves.
Real-world applications (alignment pipelines the work bears on):
- Training chat assistants and instruction-following models where preference data exists but repeated optimization epochs are common.
- Summarization and rewriting systems, the setting of the TL;DR dataset used here, where over-optimization degrades output quality.
- Any deployment where preserving model diversity matters, since collapsing probabilities push the policy away from the reference distribution.
- Practically constrained fine-tuning jobs, since SquaredPO adds no new hyperparameter beyond the existing $\beta$ and therefore needs no extra tuning budget relative to DPO.
Industry relevance. The authors are affiliated with Meta and Qualcomm AI Research alongside Technion, and the method is evaluated on Meta-Llama-3-8B-Instruct using LoRA — a configuration close to common production fine-tuning practice. Prior approaches to displacement (SimPO, $\beta$-DPO, $\varepsilon$-DPO) introduce adaptive $\beta$-s heuristically or add a hyperparameter; SquaredPO derives its adaptive $\beta$-s from theory and adds none, which matters for teams that cannot afford extra hyperparameter sweeps.
Future Directions
- Establishing a sufficient condition for displacement resistance: the paper states its condition is formally proved necessary but not established as sufficient.
- Testing other functions $f$ that satisfy both desiderata — the theory applies to general $f$, but only $f(t) = (\log t)^2/2$ is instantiated experimentally.
- Extending empirical validation beyond a single dataset and model in the main text; the paper notes its main evaluation is restricted to TL;DR and Meta-Llama-3-8B-Instruct, with additional displacement experiments on another model and another dataset deferred to Appendix C.2.
- Going beyond the abstraction of optimal solutions to account for optimizer behavior and training dynamics, which the current theoretical treatment sets aside.
- Broadening comparisons beyond DPO and $\chi$PO, and beyond LoRA finetuning, both of which the authors list as limitations.
Target Audience
Researchers and graduate students working on RLHF, DPO, and preference optimization theory; machine learning engineers responsible for alignment fine-tuning who want a loss that is competitive with DPO but more robust to long training runs; and theoretically inclined readers interested in the role of $f$-divergences, convexity assumptions, and closed-form tractability in modern alignment objectives. Comfort with $f$-divergences, Bradley-Terry models, and basic convex analysis is assumed.
Authors’ abstract
DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference policy through a KL divergence penalty. Previous work showed that this approach could be further generalized: the original problem remains tractable even if the KL divergence is replaced by a family of $f$-divergence with a convex generating function $f$. Our first contribution is to show that convexity of $f$ is not essential. Instead, we identify a more general condition, referred to as DPO-inducing, that precisely characterizes when the RLHF problem remains tractable. Our next contribution is to establish a second condition on $f$ that is necessary to prevent probability displacement, a known empirical phenomenon in which the probabilities of the winner and the loser responses approach zero. We refer to any $f$ that satisfies this condition as displacement-resistant. We finally focus on a specific DPO-inducing and displacement-resistant $f$, leading to our novel SquaredPO loss. Compared to DPO, this new loss offers stronger theoretical guarantees while performing competitively in practice.