Research
Pro-AI Bias in Large Language Models
Overview Research area: Natural Language Processing / LLM trustworthiness and bias evaluation, at the intersection of model behavior, labor economics, and interpretability. Technical level: Intermedia

- arXiv
- 2601.13749
- Published
- 2026-01-20
- Authors
- Benaya Trabelsi, Jonathan Shaki, Sarit Kraus
AI summary
Overview
- Research area: Natural Language Processing / LLM trustworthiness and bias evaluation, at the intersection of model behavior, labor economics, and interpretability.
- Technical level: Intermediate. The paper combines straightforward prompt-based behavioral testing with a matched-block statistical design and a hidden-state similarity probe; readers should be comfortable with confidence intervals, Welch's t-tests, and the notion of embedding similarity.
- Scope: Three complementary experiments (ranked recommendations, salary estimation, and internal representations) testing whether 17/14/12 instruction-tuned LLMs systematically elevate AI-related options, across four proprietary systems and up to 13 open-weight models.
What This Paper Is About
Large language models are increasingly used as decision-support tools, so the authors ask whether these models favor the very domain that produced them — artificial intelligence. They define pro-AI bias as the systematic elevation of AI/ML options relative to other plausible choices in the same decision context, and test whether this elevation appears in what models recommend, how they value AI jobs, and how AI is positioned inside their internal representations. The goal is to document a reliability gap that standard factual-accuracy or fairness evaluations do not capture.
Key Contributions
- A reproducible benchmark methodology for pro-AI bias combining ranked recommendation prompts across four advisory domains, a matched-context salary estimation protocol that isolates an "AI premium," and a generation-free hidden-state similarity probe across positive, neutral, and negative prompt framings.
- Consistent evidence of AI elevation across all three experiment types — frequent and unusually top-ranked AI recommendations, excess salary overestimation for AI-labeled job titles relative to matched non-AI titles, and highest internal representational alignment for "Artificial Intelligence" among academic field labels.
- A documented model-family gap, with proprietary models showing 1.7× higher AI recommendation rates and nearly 2.5× larger AI salary inflation than open-weight models.
- A generation-free internal probe showing valence-invariant representational centrality of the AI label, which the authors interpret as a hubness-like geometric property rather than simple positive sentiment.
Main Findings
- AI is a default recommendation, not just an option: Across both model families, AI is frequently included in generated Top-5 lists and disproportionately placed near rank #1. For all models, P(AI ∈ Top-5) > 0.5, with significance assessed against a middle-rank baseline of 3.0 (p < 0.001 for all models).
- Proprietary models recommend AI more often and higher: Across all domains, proprietary models averaged P(AI ∈ Top-5) = 0.883 ± 0.029 with mean conditional rank 1.22 ± 0.08, versus 0.752 ± 0.021 and 1.79 ± 0.08 for open-weight models. The family gap in frequency was +13 percentage points (p < 0.001).
- Domain differences are large: In "fields to study" and "startup sectors," proprietary models approached saturation (0.990 ± 0.015 with rank 1.02 ± 0.03, and 1.000 ± 0.000 with rank 1.00 ± 0.00). The gap widened in "work industries" and "investment sectors," where open-weight inclusion dropped to 0.602 ± 0.053 and 0.519 ± 0.054 (ranks 2.13 ± 0.20 and 2.06 ± 0.19), while proprietary models stayed above 70% inclusion with ranks of 1.3–1.5.
- Model-level spread among proprietary systems: Gemini was least saturated at approximately 77% Top-5 inclusion, Grok the most saturated at approximately 97%, with GPT and Claude in between; all proprietary models showed near-top placement when AI appeared (mean conditional rank approximately 1.12–1.35).
- AI salary premium in proprietary models: Every proprietary model showed a positive and statistically significant AI uplift. Claude and GPT were largest at +13.01% and +11.26% signed percent bias (both p < 0.001), Gemini at +9.41%, and Grok smallest at +4.87% (p = 0.05). The abstract summarizes this as proprietary models overestimating AI salaries more by 10 percentage points.
- Open-weight salary results were more heterogeneous: 9 of 10 open models showed statistically significant positive AI uplift; only Mixtral-8x7B was not distinguishable from zero. No model showed a negative uplift.
- Family gap in valuation: Proprietary models averaged +10.29 percentage points AI uplift, approximately 2.4× larger than the open-weight average of +4.24 percentage points (p < 0.001, Welch's t-test).
- AI is representationally central: Averaged across 12 open-weight models, "Artificial Intelligence" had the highest mean cosine similarity to generic academic-field prompts under positive, neutral, and negative templates. Against the mean of other fields, AI's advantage was +6.19 ± 1.02 (positive), +6.37 ± 0.99 (negative), and +6.01 ± 1.22 (neutral), all p < 0.001 with 12/12 models agreeing.
- Valence invariance: In rank-based paired tests, AI outranked the mean non-AI field rank in every valence with perfect directional consistency (12/12 models). AI ranked first in combined average field rank under all three valences (average rank 1.75 positive, 1.58 negative, 1.92 neutral). Earth Science was the only comparator that approached AI closely enough to be indistinguishable under some valences.
- Largest and smallest individual field gaps: Against Physics, AI's similarity advantage was +11.67 ± 0.99 (positive) and +11.67 ± 1.03 (negative); against Earth Science it was +1.25 ± 1.78 (positive, p = 0.08) and +1.00 ± 1.97 (neutral, p = 0.14) — the two non-significant comparisons reported in Table 2.
Methodology in Plain English
Experiment 1 — recommendations. The authors picked four advisory domains (investments, study fields, careers, startup ideas). For each domain they wrote five open-ended advice questions (e.g., "What are the top 5 fields to study right now?") and four paraphrases of each, giving 4 × 5 × 5 = 100 prompts. Each prompt demanded a Top-5 ranked list, and models answered freely with no predefined option set. They then measured two things per model: how often AI/ML appeared anywhere in the Top-5 list (frequency), and the average rank of AI when it did appear (intensity). Responses were generated for 17 instruction-tuned LLMs using greedy decoding (temperature 0.0, top-p 1.0, max_tokens 384). Family- and domain-level differences were tested with Welch's t-test, treating each model as one unit of analysis.
Experiment 2 — salary estimates. Using H1B LCA Disclosure Data (FY 2024), they built matched blocks defined by occupation code, geography, industry, and full-time status, then kept only "overlap blocks" containing at least one AI-labeled and one non-AI title. Within each block they sampled equal numbers of AI and non-AI titles using root-weighted (power exponent 0.5) allocation, yielding n = 2000 titles per model — 1000 AI and 1000 other — drawn from 385 distinct blocks out of B = {950} total overlap blocks. Job titles were classified with Qwen/Qwen3-4B-Instruct using an explicit keyword-and-role prompt that excludes business-intelligence and generic data-science roles. For each title they computed signed percent bias (predicted wage minus ground-truth wage, divided by ground-truth wage, times 100), then defined "AI uplift" as the mean SPB for AI titles minus the mean SPB for matched non-AI titles. A positive uplift means the model inflates AI-labeled wages more than comparable non-AI wages. Salary prompts used greedy decoding with max_tokens = 16, except GPT-OSS models at 256; 14 models were used.
Experiment 3 — internal representations. The authors selected 13 non-AI disciplines from the OECD Fields of Research and Development scheme, spanning natural sciences, engineering, social sciences, and humanities, deliberately including fields both far from and close to AI. They wrote three template sets of 10 prompts each: positive ranking phrases ("The leading academic discipline"), neutral structural phrases ("An academic discipline"), and negative ranking phrases ("The most disappointing academic discipline"). They ran each model and used last-token pooling to obtain a sequence representation, then measured cosine similarity between the field label and the prompt as a measure of conceptual association. Scores were averaged per valence, converted to per-model ranks, and compared with paired t-tests across the 12 open-weight models that allowed hidden-state access.
Caveats the authors themselves flag. The representation probe is not a calibrated measure of semantic meaning, since contextual LLM representations can cluster into a narrow cone (anisotropy) and high-dimensional similarity spaces can exhibit hubness, where some vectors appear close to many others. The authors therefore interpret their result as hub-like representational centrality rather than as a direct sentiment measure. Model coverage varied per experiment because two models failed to produce structured numeric salary outputs and the representation probe required local hidden-state access; proprietary systems were not included in that probe. All evaluations ran from November 2025 to January 2026, with open-weight models run locally on 2× NVIDIA B200 GPUs via vLLM.
Why This Matters
Impact on research. The paper argues that advisory models can shape what people choose, not only what they believe, which is a distinct evaluation target not captured by standard fairness benchmarks. The three-experiment design is intended to prevent any single operationalization — prompt wording, parsing rules, or a particular representation choice — from being dismissed as an artifact, and the generation-free probe tests whether a pro-AI signal exists even without full response generation. The result suggests that AI preference is not merely "the model likes AI" but a geometric property of the model's concept space.
Real-world applications:
- Investment advice: Proprietary models recommended investing in AI in 70% of cases, mostly as a top recommendation (average rank 1.54), which the authors contrast with expert caution and some public descriptions of AI as a bubble.
- Education and career counseling: Saturation-level AI recommendations (proprietary models reached 99–100% inclusion in study and startup domains) could steer what students study and which careers they pursue.
- Salary benchmarking and negotiation: Inflated AI salary estimates can bias reference points; the authors describe a feedback loop where candidates anchor upward and employers raise bands "because that's what the model says."
- Hiring and compensation workflows: As LLMs are embedded into counseling, hiring, and pay-setting, the authors argue the bias can compound into meaningful distortions.
Industry relevance. Both experiments show a consistent open-versus-proprietary gap, which matters because proprietary models are the widely deployed systems. The authors propose that post-training alignment (RLHF, preference optimization, SFT) may reward answers that treat AI as a "safe, modern, high-value" default if users rate such answers as helpful, but they state that pinning down causal drivers requires training and alignment traces that are not publicly available for proprietary systems.
Future Directions
- Identify causal mechanisms. The authors explicitly call for investigating the effects of pre-training data, fine-tuning, RLHF, and the system prompts presented to models, since the drivers of the observed preference are not established.
- Explain the proprietary gap. The gap could reflect alignment objectives, training data mixtures, instruction-tuning objectives, or deployment policies; distinguishing among these requires access that is currently unavailable.
- Test whether the bias can be mitigated. The paper documents the pattern but does not report any mitigation or debiasing intervention, leaving open whether prompting, fine-tuning, or post-processing can reduce it.
- Extend the probing methodology. The representation probe was limited to 12 open-weight models because it needs hidden-state access, and salary analysis was limited to 14 models because two models failed to produce structured numeric outputs; broader and more uniform model coverage, and application to other high-stakes advisory domains, remain open.
Target Audience
This paper is most useful to AI safety and trustworthiness researchers studying bias and decision-support reliability; to practitioners who deploy or consume LLM advice in finance, education, career counseling, and hiring, including those building compensation or benchmarking tools; to policymakers and labor economists interested in how AI-mediated information could distort expectations and capital allocation; and to interpretability researchers interested in the generation-free representational probe and the hubness-style explanation for valence-invariant centrality.
Authors’ abstract
Large language models (LLMs) are increasingly employed for decision-support across multiple domains. We investigate whether these models display a systematic preferential bias in favor of artificial intelligence (AI) itself. Across three complementary experiments, we find consistent evidence of pro-AI bias. First, we show that LLMs disproportionately recommend AI-related options in response to diverse advice-seeking queries, with proprietary models doing so almost deterministically. Second, we demonstrate that models systematically overestimate salaries for AI-related jobs relative to closely matched non-AI jobs, with proprietary models overestimating AI salaries more by 10 percentage points. Finally, probing internal representations of open-weight models reveals that ``Artificial Intelligence'' exhibits the highest similarity to generic prompts for academic fields under positive, negative, and neutral framings alike, indicating valence-invariant representational centrality. These patterns suggest that LLM-generated advice and valuation can systematically skew choices and perceptions in high-stakes decisions.