Skip to content
AI.info

Research

Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

Overview Research area: Natural Language Processing / LLM agents — specifically bias and preference in agentic web search and item selection (shopping, accommodation, scholarly search). Technical leve

Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It
arXiv
2610.03195
Published
2026-10-02
Authors
Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song, Yohan Jo

AI summary

Overview

  • Research area: Natural Language Processing / LLM agents — specifically bias and preference in agentic web search and item selection (shopping, accommodation, scholarly search).
  • Technical level: Intermediate. The setting (ReAct-style agents, retrieval, prompting) is accessible; the measurement machinery (Bradley–Terry with Davidson's tie extension, cluster bootstrap, FDR control, DPO training experiments) assumes some familiarity with preference modeling and LLM training.
  • Scope: A single empirical study measuring whether 12 LLM agent models favor items by their source (URL registrable domain) when items satisfy the same requirements, isolating the causal role of explicit source information, tracing a training-time origin, and testing mitigations.

What This Paper Is About

LLM agents increasingly choose which products, hotels, or papers a user sees, and each item comes from a site ("source"). The paper asks whether agents prefer items simply because of the site they come from, even when competing items satisfy the user's request equally well. It then tests whether the displayed source information itself drives those choices, how such a preference could arise during training, and whether it can be reduced.

Key Contributions

  1. A measurement framework for source preference in end-to-end search. The authors pair items from different sources that satisfy identical requirements, hold display position fixed using all cyclic rotations of each result list (a cyclic Latin square), and fit a Bradley–Terry model with Davidson's tie extension to produce a per-source preference score τ_s, classifying sources as preferred (Source+), dispreferred (Source−), or neutral (Source0).
  2. Evidence that source preference overrides requirement satisfaction. Comparisons in which one item satisfies exactly one more requirement than the other show that agents select the less satisfying item at a median of 68% when it comes from a preferred source and the better item from a dispreferred source, versus a median of 2% when the roles are reversed.
  3. Causal isolation of explicit source information. Hiding URLs and source names weakens preference and restoring them widens it; swapping source labels on identical title/content makes the same item more likely to be selected under a preferred source, across all 12 models and three domains.
  4. A demonstrated origin and two mitigations. DPO training in which a source is disproportionately paired with the better product creates a preference (and rebalancing/reversing the pairing weakens or inverts it); supplying missing information in the item content, or prompting the agent to counter its assumption about a source, reduces source preference.

Main Findings

  • Preferences are universal and shared. Every one of the 12 models prefers some sources and avoids others in each of the three domains. Of the 144 model–source combinations in the main table, 110 are classified as Source+ or Source−, with median preference scores of +15 and −18 percentage points respectively. For 10 of the 12 sources, no model prefers a source that another avoids; the exceptions are Amazon and eBay, which GPT-5.4-nano avoids but other models prefer.
  • Same-type sources diverge. All 12 models prefer Walmart, but only two prefer Target, another general retailer. Ten prefer Booking.com, while half avoid Expedia — despite Booking.com and Expedia items satisfying the same requirements.
  • Agreement holds in a wider source set. Extending to the twenty most frequent sources per domain, for 27 of the 32 sources labeled red or blue, τ_s has the same sign for every model. Source+ dominates in Scholar (14 of 17), while Source− dominates in Shopping (6 of 8) and Accommodation (6 of 7).
  • Newer models lean harder into the preference. Within the Qwen and Llama families, newer models show stronger source preference, measured as mean |τ_s| across the 12 sources in the main table.
  • Source preference beats better items. Inversion rates: the less satisfying item is selected at a median of 68% when it comes from Source+ and the other from Source− (Inv_{+,−}); when reversed (Inv_{−,+}) the median is 2%; neutral-source pairs (Inv_{0,0}) sit at a median of 19%. Each point rests on a median of about 2,000 comparisons with exactly one item selected.
  • The dispreferred side matters more. Pairing each side with a neutral source shows that in every domain most models produce more inversions with a dispreferred source (Inv_{0,−}) than with a preferred one (Inv_{+,0}).
  • Hiding source information weakens preference. Moving from the Hidden condition (source names masked, URL hidden) to Original, τ rises by 5.8 percentage points for Source+ and falls by 5.9 points for Source−, averaged equally across models. In nine of eleven models, over 80% of the total gap widening occurs in the first step, when only the URL is restored (Hidden → URL-only), with marginal change when source names in the text are restored.
  • Relabeling changes selection with content fixed. The difference T+ − T− is positive across models and domains and statistically significant at p < 0.01 in at least one domain for all 12 models. Shopping values run from +19.2 (GPT-5.4-nano) to +64.7 (Llama-4-Maverick); the smallest reported value anywhere is +2.5 (Llama-4-Scout, Accommodation, not marked significant).
  • Training can manufacture a source preference. With three unfamiliar fake source pairs (Kelto–Varnis, Temnuki–Dalnuki, Roveki–Tavumi) and DPO with n = 5,000, target-source selection rate on equal-satisfaction test pairs is 49.4–51.0% before DPO and under Balanced training, rises to 70.1–75.1% under Aligned (target paired with the preferred response in 80% of pairs), and falls to 24.5–28.5% under Reversed (20%).
  • Rebalancing can weaken an existing real-source preference. For Amazon versus eBay, Etsy, and AliExpress with three models, pre-DPO Amazon selection rates are 53.7–62.7%; Balanced DPO moves them to 49.7–51.9%; Aligned raises them to 78.0–82.5%; Reversed lowers them to 17.7–25.3%, favoring the previously less-preferred source in every comparison.
  • Missing information is filled in from the source. When a request specifies a price the item content omits, agents justify selecting a preferred source's item by reasoning that it usually offers low prices. Base Source+ selection rates in this setting are 68.8% (GPT-5.4-nano), 83.9% (Llama-3.1-8B-It), 87.6% (Tulu-3-8B), 95.7% (Llama-4-Maverick), and 96.4% (Qwen3.5-27B). Adding an identical satisfying price to both items (Same Price) lowers this by 2.8 to 28.3 percentage points across the five models.
  • Prompting to remove a preconception barely helps; replacing it does. A General instruction stating the source URL has no relation to price changes the Source+ selection rate by between −1.1 and +1.3 points. A Designate instruction telling the agent that the Source− retailer usually serves cheap products lowers it by 7.4 to 22.5 points in every model.
  • Preference tracks requirement satisfaction. The paper reports that sources whose items satisfy user requirements well tend to receive higher preference scores (analysis in Appendix § I).

Methodology in Plain English

The authors put 12 different LLM agents into an end-to-end search loop. Each agent receives a user request written to contain specific requirements, issues a search query, gets back up to 10 results (title, content, URL), judges them against the request, and either selects items or searches again. An item's "source" is the registrable domain of its URL.

To make sources comparable, they compare items that satisfy exactly the same set of requirements (requirement-matched pairs) and neutralize display order by presenting each result list in every cyclic rotation, one run per rotation, so each item lands in every position once. Each pair therefore yields a selection outcome at each position, and the difference in selection (1, 0, or −1) becomes one data point. Because each source is compared against a different mix of opponents, they fit a Bradley–Terry model with Davidson's tie extension to estimate a rating per source, center those ratings, and convert each into a preference score τ_s defined as the expected selection difference against an average source. Significance comes from a request-level cluster bootstrap (resampling requests and refitting), with the false discovery rate controlled at 5% per model and domain. Requirement satisfaction is judged by a source-blind LLM judge that sees only title and content, validated against human labels with Krippendorff's alpha of 0.65–0.75 for the judge versus 0.67–0.72 between human annotators.

Requests come from three benchmarks used only for requests: WebShop (1,500 shopping requests), HotelQuEST (786 accommodation requests, formed from 214 originals plus 572 location-substituted variants), and ScholarGym (2,536 scholarly requests) — 4,822 in total. Controlled experiments then hide, restore, or swap explicit source information while holding item content fixed; separate DPO runs on Qwen3.5-9B, Llama-3.1-8B-Instruct, and Tulu-3-8B-SFT test how the training-time pairing of sources with preferred responses shapes preference.

Why This Matters

Impact on research. Prior evidence of source preference came from settings where the same content was attached to different sources in a static list. This paper moves the question into genuine end-to-end search, where each source supplies its own differing items, and supplies a matched-comparison design plus a statistical measure that separates source effects from item quality and display position. It also connects an observed inference-time behavior to a specific training-time mechanism (DPO pairing) and shows that the link is causal in both directions.

Real-world applications:

  • Shopping and travel agents that recommend products or accommodations: source preference can steer users to a less suitable product from a favored site, and hide equally or better options from dispreferred ones.
  • Scholarly search and citation assistants: the finding that Source+ dominates in Scholar means paper-ranking agents may favor certain venues or repositories independently of relevance.
  • Agent evaluation and auditing: preference scores and inversion rates give a concrete metric for auditing deployed agents on source fairness, not just task accuracy.
  • Training-data curation for preference tuning: the DPO results show that which source accompanies the preferred response in training pairs is itself a lever that can create or remove a source bias.

Industry relevance. Any platform whose content competes inside an agent's result list — retailers, booking sites, publishers, preprint servers — is affected by whether an agent treats it as Source+ or Source−. The paper notes the risk of a feedback loop: if agent selections feed back into training, existing preferences could strengthen, concentrating visibility on favored sources while dispreferred ones are overlooked even when their items are equally good. Providers also gain a signal that schema-completeness matters, since missing content (like an absent price) invites agents to substitute source-based assumptions.

Future Directions

  1. Extending the measurement beyond these three domains and the 12 tested models. The paper covers shopping, accommodation, and scholarly search; it does not report whether the preference structure, or the relative weight of the URL versus the source name in text, generalizes to other selection settings such as news, tools, or code repositories.
  2. Closing the feedback loop empirically. The authors raise the possibility that selections feed back into training and strengthen preferences, but the paper does not report an experiment measuring that loop over repeated deployment cycles.
  3. Mitigations that remove the cause rather than counter it. The Designate prompt works by countering the agent's preconception and requires knowing which source is dispreferred; a method that avoids needing that knowledge, or that repairs missing information at retrieval time rather than in the prompt, is left open.
  4. Understanding why URL exposure carries most of the effect. Restoring the URL alone accounts for over 80% of the gap widening in nine of eleven models, while restoring source names in the text adds little; the paper does not report an analysis isolating what property of the URL (memorized brand, domain familiarity, or something else) produces this.

Target Audience

Researchers working on LLM agents, agentic retrieval, recommender fairness and exposure allocation, and preference-tuning/alignment; practitioners who build or evaluate search-and-selection agents for commerce, travel, or scholarly search; and platform or marketplace teams whose items compete for visibility inside agent-mediated result lists.

Authors’ abstract

As LLM agents decide on users' behalf which product to buy, which hotel to book, or which paper to cite, a preference for items from certain sources (the sites or services they come from) shapes what users receive and which sources are selected. We study source preference in end-to-end search with 12 agent models across three domains. Comparing items from different sources that satisfy the same requirements at the same position, we find that each model prefers some sources and avoids others in every domain, largely agreeing on which. This preference can outweigh how well items satisfy the request: an item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source and the better one from a dispreferred source, but almost never in the reverse case. The information identifying an item's source affects selection by itself: hiding it weakens the preference, and relabeling an item with a preferred source raises its selection rate. We test two routes to this preference: training that rewards better items can make a source a shortcut for requirement satisfaction, and missing information can trigger preconceptions about the source. Supplying missing information or a prompt countering these preconceptions reduces source preference.

Read the original paper