Research
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
Overview Research area: Natural language processing, specifically automatic evaluation metrics for machine translation (MT) and quality estimation (QE), plus minimum Bayes risk (MBR) decoding. Technic

- arXiv
- 2601.18006
- Published
- 2026-01-25
- Authors
- Lorenzo Proietti, Roman Grundkiewicz, Matt Post
AI summary
Overview
Research area: Natural language processing, specifically automatic evaluation metrics for machine translation (MT) and quality estimation (QE), plus minimum Bayes risk (MBR) decoding.
Technical level: Intermediate. The paper assumes familiarity with MT evaluation metrics and basic neural encoder architecture, but its core idea — predicting which of two translations is better, and by how much — is conceptually simple.
Scope: The paper introduces PEAR, a supervised QE metric family that reframes reference-free MT evaluation as a graded pairwise comparison between two candidate translations of the same source segment, and evaluates it on the WMT24 meta-evaluation benchmark and as an MBR decoding utility.
What This Paper Is About
Almost all automatic MT metrics, including reference-free QE metrics, score one candidate translation at a time and output an absolute number. Yet the main way metrics are actually used is comparative — deciding which of two systems or two candidate translations is better. PEAR addresses this mismatch by taking a source segment and two candidate translations as a joint input and predicting a signed scalar: the sign says which candidate is preferred, and the magnitude says how strong that preference is. The goal is to show that framing evaluation as a graded pairwise comparison produces a better and less redundant metric than conventional single-candidate scoring.
Key Contributions
-
A new supervised QE formulation. PEAR (Pairwise Evaluation for Automatic Relative Scoring) is introduced as a graded pairwise relative-scoring metric family, trained on pairwise supervision built by differencing human segment-level judgments, with an added regularization term encouraging sign inversion under candidate order reversal.
-
Controlled evidence for the pairwise framing. The authors train matched single-candidate QE baselines (Single-QE and Single-QE-XL) sharing the same encoder backbones, training data, and hyperparameters, isolating the effect of the pairwise formulation from backbone capacity and data exposure.
-
A reference-anchored inference mode that avoids quadratic comparisons. By fixing one side of the pair to an anchor, PEAR produces N system scores instead of the N(N−1)/2 system-to-system comparisons required by fully pairwise evaluation, and the authors show this remains effective even when the anchor is an MT output rather than a human reference.
-
PEAR as an MBR utility function. PEAR is tested as the utility for MBR decoding on 100-best lists, including a variant that exploits antisymmetry to evaluate only one triangle of the N×N utility matrix and halve the number of forward passes.
Main Findings
-
Pairwise beats matched single-candidate QE. On the WMT24 MQM test set, averaged over English–German, English–Spanish, and Japanese–Chinese, PEAR reaches higher SPA, acc*_eq, and Avg Corr than the corresponding Single-QE baseline in every matched pairing. For example, PEAR (560M) scores 80.9 SPA / 57.9 acc*_eq / 69.4 Avg Corr versus Single-QE (560M) at 80.0 / 57.2 / 68.6, and all PEAR variants take a better statistical significance rank than their Single-QE counterpart.
-
Order-reversal behavior is largely satisfied. Single-order and bidirectional (PEAR both) variants perform nearly identically, which the authors interpret as evidence that PEAR largely satisfies sign inversion under candidate order reversal.
-
PEAR beats MT-RANKER on non-tied pairs. After removing pairs with tied human MQM assessments (yielding 49,666 English–German, 24,777 English–Spanish, and 34,045 Japanese–Chinese pairs), PEAR KD (560M) reaches an average pairwise accuracy of 68.9 versus 65.8 for MT-RANKER-XXL (5.7B), and PEAR (560M) reaches 68.0.
-
Strong WMT24 results with far fewer parameters. Among reference-free metrics, PEAR-XL both (3.5B) reaches 70.2 Avg Corr, above MetricX-24-Hybrid-QE-XL (3.7B, 69.9) and XCOMET-QE (24B, 69.5), using roughly 7× fewer parameters than XCOMET-QE. PEAR both (560M) reaches 70.1 Avg Corr, above the same 3.7B and 24B metrics while using about 7× and 40× fewer parameters respectively. CometKiwi (560M), the QE metric closest in size, scores 64.0.
-
Distilled supervision helps. With additional MQM supervision distilled from GPT-4.1-mini (KD), PEAR both,KD (560M) reaches 70.4 Avg Corr, edging out CometKiwi-XXL (70.3) while being about 20× smaller (560M vs. 10.5B).
-
Competitive with reference-based metrics despite using no references. PEAR-XL both,KD (70.5 Avg Corr) matches MetricX-24-Hybrid-Large (70.5), and PEAR both (70.1) exceeds COMET-22 (68.9) and BLEURT-20 (68.6). Note that MetricX-24-Hybrid-QE-XXL (71.4), GEMBA-ESA (71.1), and PEAR-XL ref,KD (70.8) score higher on Avg Corr, and CometKiwi-XXL has higher SPA (85.4).
-
PEAR ref does not depend on human references. When the anchor slot is filled with the human reference versus the output of six MT systems — including MSLC, described as a substantially lower-quality system (Llama3-70B has a system-level MQM score of −3.62 versus MSLC's −13.46 in the English–German MQM human evaluation) — the rank stability of PEAR ref variants is stable across anchors.
-
A less redundant evaluation signal. Segment-level Pearson correlations over pairwise difference scores show PEAR both and PEAR both,KD correlate less with other top WMT24 metrics than supervised and LLM-as-a-Judge approaches. Correlation with MetricX-24-Hybrid-QE (the XXL checkpoint) is r ≈ 0.71 on English–German, r ≈ 0.51 on English–Spanish, and r ≈ 0.26 on Japanese–Chinese. Excluding lexical and unsupervised baselines (BLEU, chrF, BERTScore), PEAR metrics are the least correlated with the rest of the metric suite.
-
Effective and cheaper MBR decoding. On 100-best lists for WMT24 En-De and En-Ja, PEAR (full) versus PEAR (sym.), which evaluates only one triangle of the N×N matrix (N = 100) and fills the rest by antisymmetry, gives nearly identical results (En-De: XCOMET-XL 0.855 vs. 0.854; En-Ja: 0.810 vs. 0.809). PEAR-based MBR improves over COMET-22 and BLEURT-20 under XCOMET-XL and MetricX-24-Hybrid-XL, with smaller gains under CometKiwi-XL.
Methodology in Plain English
Model. PEAR is a cross-encoder. The source segment and the two candidate translations are concatenated into one sequence (BOS, source, SEP, candidate A, SEP, candidate B, EOS), and the encoder processes all three together. Span masks select the tokens belonging to the source and each candidate, and masked mean pooling turns them into three span vectors. For each candidate, the model builds a source-aware representation by concatenating the candidate vector, its elementwise product with the source vector, and their absolute difference. Shared parameters project each of these into a scalar utility, and the prediction is the difference between the two utilities, scaled by a learned positive factor. Only the difference is supervised, so the individual utilities are not intended as absolute quality scores.
Backbones. PEAR uses InfoXLM Large (560M parameters) and PEAR-XL uses XLM-RoBERTa-XL (3.5B parameters). The feed-forward head has hidden sizes 512 → 256 → 128 for PEAR and 2048 → 1024 → 512 for PEAR-XL.
Training. Human absolute segment-level judgments are converted into pairwise targets by subtraction: the target for a pair is the human score of candidate A minus that of candidate B. The loss combines a Huber loss on the difference with an antisymmetry term that penalizes the squared sum of the prediction and the prediction for the reversed candidate order. Training is two-stage: first pre-training on DA and DA+SQM judgments from WMT16 to WMT23, then fine-tuning on MQM supervision from WMT20 to WMT23 plus the IndicMT Eval MQM dataset for English→Indic directions. Some models are additionally fine-tuned with MQM annotations distilled from GPT-4.1-mini using a GEMBA-MQM V2 approach on language pairs without MQM coverage; these are marked KD.
Inference. Three configurations are used: the default pairwise mode (one input order), the bidirectional variant (PEAR both, which averages the scores from both orders), and the reference-anchored mode (PEAR ref, which fixes one side of the pair to an anchor translation).
Evaluation. The primary benchmark is the WMT24 Metrics Shared Task MQM test set, evaluated with the official toolkit using Soft Pairwise Accuracy (SPA) at the system level and pairwise accuracy with tie calibration (acc*_eq) at the segment level, averaged over English–German, English–Spanish, and Japanese–Chinese. Avg Corr is the mean of these two averages. Statistical significance ranks come from the PERM-BOTH hypothesis test, following the WMT24 setup.
Why This Matters
Impact on research. The paper argues that the dominant single-candidate, absolute-score paradigm is structurally mismatched to how MT metrics are actually used, and provides controlled evidence that a pairwise formulation improves comparison quality under matched conditions. The finding that PEAR's segment-level scores correlate weakly with other top metrics suggests metric diversity that could matter for MT fine-tuning, where the authors note diverse metrics help avoid over-optimizing for a single signal.
Real-world applications:
- MT system selection and leaderboard-style comparisons, including shared task evaluations that rely heavily on metric scores.
- Ranking candidate translations within a decoding or post-editing pipeline.
- MBR decoding, where PEAR provides a utility function trained explicitly for reference-free pairwise comparison, and where the antisymmetry shortcut roughly halves the number of forward passes.
- Fine-tuning of MT models, where a metric less correlated with existing metrics may provide a complementary training or selection signal.
Industry relevance. The efficiency story is central: PEAR matches or exceeds much larger QE models and reference-based metrics while being roughly 7× to 40× smaller than the comparable ones, and its reference-anchored mode avoids the quadratic N(N−1)/2 cost of full system-to-system comparison. Both factors matter for deployment at scale, where metric cost and throughput are practical constraints.
Future Directions
-
Scaling. The largest checkpoint fine-tuned is PEAR-XL at 3.5B parameters. The authors explicitly state they do not test whether the gains attributed to the pairwise formulation persist, widen, or saturate with larger backbones, and note that scaling is a known strong driver of performance for supervised metrics.
-
From scalar scores to error spans. The pairwise setup could be extended to MQM sequence tagging that explicitly predicts error spans in each candidate translation, aligning with side-by-side MQM protocols based on comparative judgment.
-
Understanding the correlation gap. Why PEAR's segment-level scores diverge from other top metrics is left to future work; the authors state that investigating which phenomena drive this divergence is a priority.
-
Applying pairwise relative scoring elsewhere. The conclusion states that future work will focus on leveraging pairwise relative scoring in other evaluation settings beyond what is tested here.
Target Audience
MT evaluation researchers and practitioners, including those working on quality estimation metrics, WMT-style meta-evaluation, and metric development; researchers applying MBR decoding or other preference-based decoding methods; and engineers who need efficient, reference-free scoring for system selection or model fine-tuning. Readers with a background in neural MT and evaluation benchmarks will get the most out of it, though the core idea is accessible without deep architectural knowledge.
Authors’ abstract
We present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-free machine translation (MT) evaluation as a graded pairwise comparison. Given a source segment and two candidate translations, PEAR predicts the direction and magnitude of their quality difference. The metrics are trained using pairwise supervision derived from differences in human judgments, with an additional regularization term that encourages sign inversion under candidate order reversal. On the WMT24 meta-evaluation benchmark, PEAR outperforms strictly matched single-candidate QE baselines trained with the same data and backbones, isolating the benefit of the proposed pairwise formulation. Despite using substantially fewer parameters than recent large metrics, PEAR surpasses far larger QE models and reference-based metrics. Our analysis further indicates that PEAR yields a less redundant evaluation signal relative to other top metrics. Finally, we show that PEAR is an effective utility function for minimum Bayes risk (MBR) decoding, reducing pairwise scoring cost at negligible impact.