Research
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
Overview Research area: Evaluation of large language models (LLMs), specifically human preference measurement, psychometrics, and demographic fairness in NLP. Technical level: Intermediate. The framin
- arXiv
- 2603.04409
- Published
- 2026-02-03
- Authors
- Nora Petrova, Andrew Gordon, Enzo Blindow
AI summary
Overview
- Research area: Evaluation of large language models (LLMs), specifically human preference measurement, psychometrics, and demographic fairness in NLP.
- Technical level: Intermediate. The framing and results are accessible, but the statistical engine (a hierarchical Bayesian Bradley-Terry-Davidson model with post-stratification) requires some comfort with Bayesian modelling and ranking theory.
- Scope: The paper introduces HUMAINE, a framework for multidimensional, demographically stratified human preference evaluation, and reports results from 119,890 human judgments across 23,404 participants, 22 demographic strata, 28 models, and five evaluation dimensions in the US and UK.
What This Paper Is About
LLM evaluation currently splits into automated benchmarks, which measure technical capability but miss subjective interaction quality, and human preference platforms, which capture real conversations but rely on self-selected anonymous users, shallow judgments, and a single binary preference vote. The paper's goal is to close this "evaluation gap" by measuring human-AI interaction quality in a way that is multidimensional, demographically representative, and statistically rigorous. It does so by collecting naturalistic multi-turn conversations from a census-stratified participant pool and modelling preferences with a hierarchical Bayesian ranking model.
Key Contributions
- The HUMAINE framework: a methodology for human-centric AI evaluation that explicitly targets three validity threats in prior work, namely sampling bias, shallow assessment depth, and single-metric reductionism.
- A large-scale, demographically stratified dataset: 119,890 multidimensional human judgments from 23,404 participants comparing 28 models, plus structured metadata characterising conversational dynamics, task properties, and interaction outcomes. The dataset is released publicly.
- Empirical insights into preference heterogeneity: analysis showing how model rankings shift across demographic groups and across evaluation dimensions, with implications for context-appropriate model selection.
- A living evaluation framework: a regularly updated public leaderboard (released on Hugging Face Spaces) that tracks state-of-the-art models as new ones are released.
Main Findings
-
Clear top-ranked model, with quantified confidence:
google/gemini-2.5-proranks first overall, with a 95.6% posterior probability of being the best model. A distinct gap separates it from the runner-up,deepseek/deepseek-chat-v3-0324, which in turn leads a competitive tier includingmistralai/magistral-medium-2506,x-ai/grok-4, andx-ai/grok-3with closely overlapping credible intervals. Lower-ranked models are largely statistically indistinguishable. -
Age is the dominant axis of preference disagreement: A model's average rank shifts by ±2.8 ranks across age cohorts, versus ±1.3 for ethnicity and ±1.5 for political affiliation. For example,
mistralai/magistral-medium-2506ranks 1st in the US and 2nd in the UK among 18-34 users but falls to 5th (US) and 10th (UK) among users aged 55+, whilegoogle/gemini-2.5-proimproves with age and tops older cohorts in both regions. -
Older users are measurably less decisive: Tie rates rise from 9.7% for the 18-34 cohort to 12.5% for users aged 55+, a 29% increase in indecisiveness. On Core Task Performance & Reasoning specifically, tie rates climb from 32% (18-34) to 39% (55+), whereas all age groups remain similarly decisive about Communication Style & Presentation (17-20% ties).
-
Rankings change with the evaluation lens: While
google/gemini-2.5-proholds the top position across all dimensions, other models move substantially.x-ai/grok-3ranks 2nd on Core Task Performance & Reasoning but 8th on both Communication Style & Presentation and Interaction Fluidity & Adaptiveness.mistralai/magistral-medium-2506ranks 2nd on Interaction Fluidity & Adaptiveness but 7th on Core Task Performance & Reasoning and 12th on Trust, Ethics & Safety. -
Dimensions differ enormously in discriminative power: Trust, Ethics & Safety showed the highest tie rate at 65%, making it the least discriminative dimension, while Overall Winner was the most decisive at 10% ties. The paper argues this means holistic judgments give a strong signal in open conversations, while nuanced safety attributes need specialised scenarios to be assessed meaningfully.
-
Open-ended conversations spanned real-world use: Post-hoc LLM classification found the majority of conversations involved information seeking (71.5%), followed by personal advice (10.5%), project planning (2.7%), technical assistance (2.4%), and decision support (2.2%). Conversations covered 41 distinct domains, most commonly health/medical (12.9%), sports (8.8%), technology (8.1%), cooking/food (7.6%), and creative arts (7.5%). Mean task complexity was 3.54 (median 4.0), mean goal achievement 4.32 (median 4.0, with 92.6% scoring 4-5), and mean user engagement 3.30 (median 3.0).
-
Disparity with technical benchmarks: The paper notes that
google/gemini-2.5-procurrently ranks a modest 13th on HELM, illustrating the gap between technical accuracy ranking and human preference ranking.
Methodology in Plain English
Participants were recruited through Prolific and paid £9/hr, then stratified into 22 demographic groups defined by geography (US/UK), age (18-34, 35-54, 55+), ethnicity (Asian, Black, White, Other in the UK; Hispanic, Asian, African American, White in the US), and political affiliation (Democrat, Republican, Independent in the US; Conservative, Labour, Liberal Democrats, Greens, Reform UK in the UK).
Each participant chose their own conversation topic and interacted with two anonymised models side-by-side. A single input field sent identical messages to both models at once, so both models handled the exact same conversational context rather than diverging trajectories. A minimum of 3 conversational turns was required (median conversation length was 6 turns), after which participants chose which model was better or declared a tie across the five dimensions. Model pairings were chosen adaptively by a TrueSkill-based algorithm that selects matchups with the most uncertain outcomes to speed up ranking convergence. Each stratum ran as its own TrueSkill tournament, with 1,848 to 2,636 comparisons per stratum, and participants could qualify for multiple tournaments based on their demographic profile; analysis pools all comparisons and disentangles the mixed effects statistically. A gpt-4o-mini judge monitored conversations in real time for low-effort input and removed participants after three warnings, affecting less than 1.6% of the sample.
The analysis uses a hierarchical Bayesian Bradley-Terry-Davidson (BTD) model. It learns a global skill parameter for each model-metric pair, adds hierarchical, centred demographic adjustments for age, ethnicity, and politics, and includes a per-metric tie propensity. Adjustments are scaled so that three demographic axes remain comparable to one, and heterogeneity parameters (τ) quantify how much preferences vary. Results are post-stratified to census weights so the final leaderboard reflects the broader US and UK populations. Separately, a gpt-4.1 judge performed post-hoc analysis of all transcripts to produce explanatory metadata; the paper states this analysis was strictly separated from ranking and never influenced human preference scores.
Why This Matters
Impact on research. The paper reframes LLM evaluation from "which model is best?" to "best for what and for whom?", arguing that aggregate leaderboards hide performance trade-offs, mask demographic blind spots, and misrepresent how useful different metrics actually are. It also provides a reusable methodological template, combining psychometric ranking theory with census post-stratification, and releases the dataset, an interactive leaderboard, and open-source code.
Real-world applications:
- Model selection by organisations: Teams can pick models whose dimensional strengths match their use case, rather than defaulting to a single leaderboard position.
- Product and UX decisions for consumer AI: Knowing that younger and older users weight and distinguish dimensions differently informs interface and default-behaviour choices.
- Equity and representation auditing: Stratified evaluation exposes performance gaps that anonymous, self-selected samples typically hide, supporting fairer deployment across populations.
- Safety and trust assessment design: The 65% tie rate on Trust, Ethics & Safety signals that generic conversations are a poor instrument for safety evaluation, pointing to a need for targeted scenario-based test suites.
Industry relevance. The findings suggest development teams tuned on narrow, tech-savvy feedback risk preference optimisation loops that exclude broader populations, undermining both adoption and equitable performance. The paper also flags market-level risks from unrepresentative evaluation, noting that undisclosed private testing and evaluation gaming can distort rankings independently of true model quality.
Future Directions
- Geographic and demographic expansion. The current version covers only US and UK populations due to budget constraints and census data availability for post-stratification. The framework is designed to be extensible to additional geographies, languages, and demographic dimensions such as gender, education, and socioeconomic status.
- Longer and more controlled interactions. The focus on short multi-turn conversations cannot capture long-term phenomena like persona consistency or performance degradation over extended dialogues, and the open-ended design means task complexity was not controlled.
- Multimodal evaluation. The framework is currently text-only, which the authors say assesses only a fraction of the capabilities of increasingly multimodal state-of-the-art models. Designing tasks that evaluate reasoning and coherence across modalities is described as a significant research challenge.
- Targeted evaluation suites for nuanced qualities. Because Trust, Ethics & Safety showed a 65% tie rate, the authors call for specialised scenarios, such as sensitive topics, ethical boundary navigation, and high-stakes advice requests, that create the context needed for users to make discriminative judgments. They also note that dimensions like creativity, humour, and empathy may be significant preference drivers that the five current dimensions may not exhaust.
Target Audience
This paper is most useful for LLM evaluation researchers and benchmark designers, applied ML and policy teams responsible for responsible-AI and fairness assessments, product and research leads at AI companies choosing or tuning models for diverse user bases, and HCI or psychometrics researchers interested in applying measurement theory to AI evaluation. Readers need only a general familiarity with how model leaderboards work; the Bayesian methodology is presented in plain language in the main text with full mathematical detail relegated to the appendix.
Authors’ abstract
The evaluation of large language models faces significant challenges. Technical benchmarks often lack real-world relevance, while existing human preference evaluations suffer from unrepresentative sampling, superficial assessment depth, and single-metric reductionism. To address these issues, we introduce HUMAINE, a framework for multidimensional, demographically aware measurement of human-AI interaction. We collected multi-turn, naturalistic conversations from 23,404 participants that were stratified across 22 demographic groups, both in the US and UK, to evaluate 28 state-of-the-art models across five human-centric dimensions. We use a hierarchical Bayesian Bradley-Terry-Davidson (BTD) model, with post-stratification to census data, and our analysis reveals three key insights. \textbf{(1)} We establish a clear performance hierarchy where \texttt{google/gemini-2.5-pro} ranks first overall, with a 95.6\% posterior probability of being the top-ranked model. \textbf{(2)} We uncover significant preference heterogeneity, with user age emerging as the primary demographic axis of disagreement; a model's perceived rank can shift substantially across age groups, exposing failures in generalisation that unrepresentative samples typically mask. \textbf{(3)} We quantify the vast difference in discriminative power across evaluation dimensions, with ambiguous qualities like \textit{Trust, Ethics \& Safety} showing a 65\% tie rate, in stark contrast to the decisive 10\% tie rate for \textit{Overall Winner}. Our work emphasises the need for a more multidimensional, demographically aware perspective in LLM evaluation. We release our complete dataset, interactive leaderboard, and open-source framework.