Research
Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
Overview Research area: Natural Language Generation (NLG) evaluation, specifically "LLM-as-a-judge" frameworks and the statistical reliability of automated raters. Technical level: Intermediate. Famil

- arXiv
- 2510.27106
- Published
- 2025-10-31
- Authors
- Rajarshi Haldar, Julia Hockenmaier
AI summary
Overview
- Research area: Natural Language Generation (NLG) evaluation, specifically "LLM-as-a-judge" frameworks and the statistical reliability of automated raters.
- Technical level: Intermediate. Familiarity with NLG benchmarks (summarization, dialogue) and agreement statistics helps, but the core argument is explained in plain terms.
- Scope: A measurement study of how consistently three open-weight LLM judges rate the same items across repeated runs on three benchmarks (SummaC, SummEval, MT-Bench), and how that self-inconsistency relates to their agreement with human judges.
What This Paper Is About
LLM-as-a-judge has become a common way to evaluate generated text, and these judges are usually validated by comparing their scores against human scores taken as the gold standard. The authors point out that almost no study reports a judge's agreement with itself across repeated runs under identical prompts and hyperparameters — a property they call self-reliability or intra-rater reliability. Their goal is to quantify that self-inconsistency across several NLG tasks and benchmarks, and to test whether any practical workaround (majority voting over runs, or disabling sampling) preserves the judge's usefulness.
Key Contributions
- Shows that ratings output by LLM judges have low agreement across multiple runs using the same prompt and hyperparameters.
- Shows that turning off sampling so the model always emits the same rating hurts performance as measured by agreement with human judgment, creating a trade-off between self-reliability and accuracy.
- Demonstrates that the phenomenon persists across multiple NLG tasks and benchmarks — binary factual-consistency labeling (SummaC), Likert-scale summarization scoring on four metrics (SummEval), and pairwise conversation ranking (MT-Bench).
- Provides recommendations for more robust NLG evaluation, including reporting intra-rater reliability, aggregating across runs, and collecting self-reliability data on human annotators.
Main Findings
- Self-reliability on SummaC is low and model-dependent. Measured by Krippendorff's Alpha over three runs: Llama 3.1 = 0.3263, DeepSeek-R1 = 0.6278, Qwen-3 = 0.7883. Only the newest/largest model approaches the commonly accepted 0.8 threshold for good agreement.
- Reliability depends heavily on the evaluation metric in SummEval. DeepSeek-R1 and Qwen 3 show high self-reliability on Coherence and Consistency but very low self-reliability on Fluency; Fluency is also the only metric where Llama performs best. The paper does not report explicit numeric self-reliability values for SummEval in the text.
- Pairwise ranking (MT-Bench) is the least stable task. Self-reliability was 0.265 for Llama 3.1, 0.507 for DeepSeek-R1, and 0.563 for Qwen 3 across three runs — all far below the 0.8 threshold. Qwen 3 gave the same judgment on all three runs for only 61.3% of cases.
- Majority voting over runs improves agreement with humans. Balanced accuracy against human judgments on SummaC: Llama 3.1 61.4 (vs. 59.1 ± 2.06 for a single run), DeepSeek-R1 72.3 (vs. 69.8 ± 0.50), Qwen 3 80.6 (vs. 79.4 ± 0.32). For DeepSeek-R1 and Qwen 3, the majority vote beats the best single run.
- Disabling sampling degrades performance. Running with no sampling gave balanced accuracies of 58.4 (Llama 3.1), 69.3 (DeepSeek-R1), and 79.2 (Qwen 3) — lower than both the single-run mean and the majority vote for every model.
- Prompted LLM judges can rival fine-tuned baselines on SummaC. Qwen 3 significantly outperformed the SummaC_ZS and SummaC_CONV baselines overall, while DeepSeek-R1 was competitive; the baselines were fine-tuned on the task whereas the judges were prompted off-the-shelf. DeepSeek-R1 performed best on XSumFaith, and SummaC_CONV performed best on FactCC (unsurprising, since it was fine-tuned on that dataset).
- Accuracy inflates agreement compared to chance-corrected metrics. On MT-Bench, human-vs-human agreement was 0.827 by accuracy but only 0.478 by Krippendorff's Alpha — below the 0.8 threshold even for humans.
- GPT-4 is the strongest LLM judge on MT-Bench, but still far below human agreement. Accuracy / Krippendorff's Alpha: GPT-4 0.671 / 0.396, Qwen 3 0.719 / 0.426, DeepSeek-R1 0.668 / 0.385, Llama 3.1 0.556 / 0.239.
- Human judges disagree with each other on SummEval, and judge type matters. Experts showed the highest agreement on consistency (0.798), moderate on fluency (0.588), and lowest on relevance (0.398). Crowdworkers ("Turkers") were uniformly lower across all metrics (0.48–0.51). Expert-vs-Turker agreement was drastically lower (maximum 0.247 for relevance). Experts and LLMs showed modest agreement, highest on consistency and coherence, dropping or turning negative for subjective metrics such as fluency; even the best observed agreement (0.726 for consistency) stayed below the accepted substitution threshold.
Methodology in Plain English
The authors took three existing benchmarks covering different flavors of NLG evaluation and ran three open-weight LLMs as judges on each:
- SummaC — label a summary as consistent or inconsistent with its source article (binary). It unifies six datasets (CoGenSumm, XSumFaith, Polytope, FactCC, SummEval, FRANK).
- SummEval — 1,700 examples, each with model-generated summaries rated on a 1–5 scale for coherence, consistency, fluency, and relevance, with scores from both 3 expert and 5 crowd-sourced annotators.
- MT-Bench — multi-turn conversations where a rater picks model_a, model_b, or tie; 80 questions with 30 examples each (2.4k examples total), plus a filtered subset of 761 examples that have two or more human ratings so agreement among humans can be computed.
The judge models were Llama-3.1-70B-Instruct, DeepSeek-R1-Distill-Qwen, and Qwen3-32B. Each judge ran the same items three independent times with identical prompts and settings, so any variation in scores is attributable to the model's own sampling rather than to changing conditions. The authors note they also tried additional runs (up to 10 in initial experiments, and up to 5 when repeating the SummaC experiments) and found no significant effect of the number of runs on self-reliability, so they settled on 3.
To measure agreement they used Krippendorff's Alpha rather than accuracy or correlation, because those metrics do not correct for chance agreement. The authors illustrate the problem: with binary labels appearing 95%/5% of the time, two raters would be expected to agree 90.5% of the time by chance (0.95 × 0.95 = 0.9025 and 0.05 × 0.05 = 0.0025), so an observed agreement of 90% is actually worse than chance. The distance function was adapted per benchmark: nominal distance for SummaC's categorical labels, and ordinal distance for SummEval and MT-Bench, where adjacent ratings are closer than distant ones. Balanced accuracy (the mean of sensitivity and specificity) was also reported for SummaC to handle class imbalance, and plain accuracy was kept for MT-Bench to match the original benchmark. For SummaC, the three runs were combined by simple majority vote before comparing to the human label. Experiments ran on a 4xA100 GPU server using the transformers library, with default recommended settings per model (temperature 0.6 with top_p 0.9 for Llama 3.1 and DeepSeek-R1; temperature 0.6 with top_p 0.95 for Qwen 3).
Why This Matters
Impact on research. Single-run LLM judgments are routinely reported as if they were stable measurements. If a judge's own ratings vary substantially between identical runs, then reported correlations with human judgment (and comparisons between competing judges or evaluation frameworks) are built on a noisy foundation. The paper also raises the uncomfortable possibility that the human "gold standard" may carry the same self-consistency problem — most benchmarks contain only one human judgment per example, making it impossible to check.
- Real-world application — content summarization: systems that grade or filter auto-generated summaries by quality may disagree with themselves from one pass to the next.
- Real-world application — automated journalism and translation: quality gates that rely on an LLM judge can wave through or reject the same output depending on the run.
- Real-world application — customer service chatbots: choosing between competing assistants by LLM preference ranking is unstable on multi-turn dialogue, the exact setting MT-Bench tests.
- Real-world application — model leaderboards: crowdsourced leaderboards such as Chatbot Arena already face equity concerns; unstable automated judges add another layer of noise.
Industry relevance. Any organization using LLM judges to build evaluation pipelines, select between model checkpoints, or augment human review needs a defensible answer to "how repeatable is this score?" The paper's practical guidance — aggregate over multiple runs rather than disabling sampling, and report chance-corrected agreement metrics instead of raw accuracy — is directly actionable for evaluation teams.
Future Directions
- Measuring human self-reliability. Collecting repeated judgments from the same human annotators would establish an upper bound on the self-reliability reasonably expected of LLM judges, and would clarify how much training or expertise affects it.
- Improving self-reliability without losing accuracy. The paper identifies a trade-off between consistent ratings and agreement with humans but does not resolve it; methods that raise both are an open problem.
- Reasoning traces and model internals. The authors did not explore the relationship between a judge model's reasoning traces and its self-reliability, nor whether probing specific layers could explain why conflicting runs diverge.
- Prompt structure and fine-tuning. The paper did not investigate how prompt design or supervised fine-tuning affect judge reliability, and notes that few-shot and chain-of-thought prompting produced no reliability or agreement gains in its own additional experiments.
- More objective tasks and proprietary models. Quantifying how much of the inconsistency comes from the inherent subjectivity of summarization and dialogue (by comparison with more objective tasks) remains open, as does testing the phenomenon on newer proprietary models such as GPT-5 and Claude-4.
Target Audience
Researchers and practitioners who rely on LLM-as-a-judge evaluations or human annotation for generated text — NLG and NLP evaluation researchers, benchmark and leaderboard maintainers, and machine learning engineers building automated quality-assessment pipelines. It is also relevant to anyone designing annotation studies, since it argues that intra-rater reliability should be reported alongside inter-rater reliability for both machine and human judges.
Authors’ abstract
As Natural Language Generation (NLG) continues to be widely adopted, properly assessing it has become quite difficult. Lately, using large language models (LLMs) for evaluating these generations has gained traction, as they tend to align more closely with human preferences than conventional n-gram or embedding-based metrics. In our experiments, we show that LLM judges have low intra-rater reliability in their assigned scores across different runs. This variance makes their ratings inconsistent, almost arbitrary in the worst case, making it difficult to measure how good their judgments actually are. We quantify this inconsistency across different NLG tasks and benchmarks and see if judicious use of LLM judges can still be useful following proper guidelines.