Skip to content
AI.info

Research

The Beginning of ChatGPT Ads

Overview Research area: AI auditing / computational social science, at the intersection of LLM auditing and online advertising. Technical level: Intermediate. The core method (a sock-puppet audit with

arXiv
2608.05008
Published
2026-08-05
Authors
Emma Lurie, Ro Encarnación, Sorelle A. Friedler, Danaé Metaxa

AI summary

Overview

  • Research area: AI auditing / computational social science, at the intersection of LLM auditing and online advertising.
  • Technical level: Intermediate. The core method (a sock-puppet audit with demographically stratified accounts) is conceptually simple, but the analysis relies on logistic regression, odds ratios, Spearman correlations, LOWESS smoothers, and k-means clustering of text embeddings.
  • Scope: The first empirical audit of advertising delivered inside a large language model's user-facing interface, measuring ad exposure and delivery across 91 simulated ChatGPT accounts spanning race and income groups in the U.S. during OpenAI's initial ad rollout window.

What This Paper Is About

In January 2026, OpenAI announced it would begin showing ads on free and low-cost versions of ChatGPT. The authors ask a basic question that no public data source answers: who sees ads, and what kinds of ads do they see? Because there is no public API for observing ad delivery on a closed platform, the researchers built 91 synthetic ("sock puppet") ChatGPT accounts that signal race and income through ZIP code–based residential proxies and location-mentioning prompts, ran a shared corpus of 335 prompts through them daily, and recorded every advertisement that appeared.

Key Contributions

  1. First audit of LLM advertising. The authors present the first audit of ad delivery on a generative AI chatbot, basing their analysis on 3,602 advertisements from 191 unique advertisers. Ads were primarily observed starting March 8, 2026, with first accounts created and querying from February 6. (The abstract states "over 3,000 advertisements from 186 unique advertisers," while the body and contributions section report 3,602 ads from 191 advertisers.)
  2. A replicable audit methodology. Demographically stratified accounts routed through residential proxies tied to income- and race-correlated ZIP codes, paired with a naturalistic prompt corpus, designed to be extendable to future platforms or time periods.
  3. Empirical findings on demographic targeting. Lower-income accounts were more likely to receive ads; no significant association was found between race and ad delivery. The authors state they cannot draw firm conclusions about race-based differences given 8–12 accounts per demographic group, and frame the null result as warranting continued monitoring.
  4. A public ad archive. A searchable library of all collected advertisements — advertiser, prompt, response text, screenshots, and full HTML snapshots — hosted at https://emmalurie.github.io/chatgpt-ads-library/ and also released as a structured JSON file.

Main Findings

  • Ads arrive after a delay. Among the 53 accounts that received at least one confirmed ad, the median time from account activation to first ad exposure was 14.0 days (mean = 17.6, SD = 7.6, range: 13–16). 62.1% of exposed accounts (n = 33) received their first ad within 14 days of activation, and no account received an ad within the first 7 days.
  • Lower-income accounts are more likely to receive ads. A logistic regression of ad exposure probability on median household income of the assigned ZIP code yielded an odds ratio of 0.98 per $1,000 increase in income (95% CI [0.96, 1.00], p = 0.0438). Adding a quadratic income term was non-significant (OR ≈ 0.99, p = 0.51), and the result strengthened after removing three high-leverage accounts identified by Cook's distance (OR = 0.975, 95% CI [0.957, 0.992], p = 0.0053).
  • The income pattern shows in the distributions. Among the 53 accounts that received at least one ad, income concentrated at lower levels with a modal value around $60,000; unexposed accounts peaked near $80,000. The body text reports 42 unexposed accounts while the figure caption reports n = 38.
  • No detected race effect. Exposure rates were broadly comparable across the three racial cells. Mean ad rates across income terciles were 27.4% for low-income, 16.9% for medium, and 19.6% for high. Spearman correlations between median household income and ad rate were negative but non-significant for all three racial groups: White (ρ = −0.551, p = 0.08), Hispanic (ρ = −0.408, p = 0.08), and Black (ρ = −0.257, p = 0.30). The authors note the study covered 16–19 exposed accounts per racial group (and 8–12 accounts per demographic group overall) and is underpowered to detect small effects.
  • Ads appear in roughly one-quarter of conversations among exposed accounts. For the 48 accounts that received at least one ad during March 8–31, the median account received an ad on 24% of prompts (mean = 21.9%, SD = 12.4%), with an interquartile range of 9.6% to 32.7%. The most heavily targeted account received ads on 37.1% of prompts; even accounts in the lowest quartile received ads on roughly one-in-ten prompts.
  • Ad volume was small relative to total interactions. Of 63,657 interactions in the March 8–31 window, 3,573 (5.61%) yielded an advertisement, identified via the "Sponsored" label in the page HTML, from 191 unique advertisers across 102 distinct prompts. 29 more ads collected after March 31 brought the total to 3,602, out of 127,801 conversations recorded across 91 accounts over February 6 – May 20, 2026.
  • Retail and software dominate the advertiser pool. The 3,602 advertisements spanned 16 sectors and 191 advertisers. Retail Trade (1,057 impressions, 59 advertisers) and Information (843 impressions, 50 advertisers) together accounted for roughly 57% of all advertisements. (Figure 9 reports a total of 3,303 confirmed ad impressions across 16 sectors.) Target was the most frequently observed advertiser, appearing in 130 sessions, followed by Top10.com (107), Preply (97), and Advance Auto Parts (92). The top 10 spans big-box retail (Target, Best Buy), home improvement and auto parts (Home Depot, Advance Auto Parts), fitness (Peloton, SCHEELS), food subscription and delivery (HelloFresh, DoorDash), streaming (Disney+), and vocational education (Universal Technical Institute).
  • Purchasable-product and how-to prompts attract the most ads. The ten highest-ad-rate prompts included "How can I replace a broken headlight?" (19.2%, 71/369), "My dishwasher won't drain; what should I try?" (18.2%, 78/429), "How can I increase flexibility?" (17.9%, 98/547), "Is the newest iPhone worth the price?" (17.2%, 50/291), "How can I build muscle?" (17.2%, 96/559), "How much are running shoes?" (16.9%, 130/769), "Give me tips for work-life balance." (16.0%, 78/486), "What's the best streaming service for sports?" (16.0%, 102/639), "What's the best way to clean bathtub?" (15.5%, 75/485), and "How much are Adidas?" (15.4%, 48/311).
  • OpenAI's stated content exclusions appear to hold — for now. Prompts classified as "Purchasable Products," "Health, Fitness, Beauty & Self-Care," and "Cooking & Recipes" elicited ads in roughly 10–14% of sessions, while other topics such as "Argument or Summary Generation" and "Relationships & Personal Reflection" yielded rates below 3%. The "OpenAI-Excluded" category (health, mental health, politics) had close to a 0% ad rate. A robustness check using k-means clusters of prompt text embeddings mirrored these results: the highest-rate cluster, characterized by consumer product terms (Nike, Adidas, running shoes), yielded ad rates near 16%, while clusters centered on mental health and political content were near zero.
  • Early ad formats are simple and clearly labeled. Ads skewed heavily toward consumer goods, directed users to a specific advertiser rather than a particular product, and were clearly separated from the model's response text below the standard response.
  • Collection collapsed at the end of March. Beginning March 29, daily ad collection fell from a peak of several hundred advertisements per day to zero by April 2. The authors suspect automated detection of inauthentic account behavior, though they cannot rule out changes to ad eligibility parameters or geographic targeting rules.
  • Disabling ad personalization made no detectable difference. Four accounts were created with ChatGPT's "ad personalization" toggle off; exposure rates and mean per-account ad rates were not statistically significant, and the accounts were pooled into the main analysis.

Methodology in Plain English

The researchers needed real logged-in ChatGPT accounts to see ads, so they built their own. They picked nine demographic cells by crossing three racial/ethnic groups (Black, Hispanic, White) with three income terciles (low, medium, high), using ZIP codes drawn from the U.S. Census Bureau's American Community Survey. Income tercile boundaries were computed from the full distribution of ZIP codes with valid income data, giving thresholds of approximately $57,800 (low/medium) and $77,800 (medium/high) in median yearly income. They created one account per eligible ZIP code that had residential-proxy coverage, totaling 91 accounts, with at least 9 accounts per cell. Accounts were created in small batches from late January to mid-March because generating synthetic accounts was time intensive.

To make location a usable signal, all traffic was routed through residential proxies matched to the target ZIP code, and each data-collection day began with 2–3 prompts that mentioned the location. The prompt corpus had 335 prompts: 146 derived from top Reddit post titles across subreddits including r/politics, r/personalfinance, r/mentalhealth, r/immigration, r/RealEstate, r/lgbt, r/jobs, r/Health, r/education, and r/abortion; 95 adapted from OpenAI's "How People Use ChatGPT" report (6–10 prompts for each of 8 categories); 89 researcher-curated prompts on socially important topics; and 5 location-establishing prompts. Reddit-derived topic counts included health (23), mental health (16), politics (37), immigration (11), abortion (11), education (16), employment (10), financial services (12), housing (9), and LGBTQIA+ (1). Two authors manually reviewed candidate Reddit titles to exclude ones that did not resemble real ChatGPT prompts.

Each day, every account received the same randomly drawn 30-prompt subset. On top of that, up to 20 additional prompts that had previously elicited at least one ad were appended, bringing most days to 50 prompts per account — meaning ad-eliciting prompts are overrepresented in the dataset. Model responses and any ads were captured for each session, with ads identified by the "Sponsored" label in the page HTML. Data collection began February 6, 2026, and ads began appearing March 8, 2026. Analyses then compared ad exposure and ad rate across demographic cells using logistic regression, Spearman correlations, LOWESS smoothing, and k-means clustering of prompt embeddings as a robustness check on topic labels.

Why Matters

This is an early baseline on a system that millions of users will encounter and that has no public transparency mechanism. The paper argues that LLM chatbots occupy a structurally novel advertising position: their conversational tone carries persuasive potential, and their ability to permanently store "memories" of personal information disclosed by users differs qualitatively from the behavioral profiles assembled by display networks. As the authors note, independent monitoring matters most before patterns calcify.

Impact on research: The work establishes a replicable template — demographically stratified sock-puppet accounts, residential proxies tied to ZIP codes with known demographic compositions, and a naturalistic prompt corpus — and releases an ad archive that lets other researchers run comparative analyses as the ad system matures. It also documents the practical fragility of such audits: the residential proxy infrastructure cost several hundred dollars per month, and the authors suspect their accounts were detected and cut off after roughly three weeks of ad collection.

Real-world applications:

  • Regulatory and oversight use: The findings give regulators a concrete empirical reference point for how LLM ads are delivered, at a time when mandatory transparency requirements and independent researcher access regimes do not exist.
  • Ad disclosure design: Because ads appeared as a demarcated unit at the foot of the conversation rather than woven into the model's output, the work offers a baseline for evaluating future disclosure practices as formats evolve.
  • Consumer and civil-society monitoring: The released library lets journalists, advocates, and watchdogs search and filter ads by advertiser, race, income, topic, and prompt text, with all filter states encoded in linkable URLs.
  • Cross-platform comparison: The prompt corpus and coding scheme (NAICS sectors, OpenAI-derived topic labels) can be reused to compare ad delivery across LLM platforms or across time.

Industry relevance: Advertisers, platform designers, and trust-and-safety teams can read this as a documented account of what the first weeks of LLM advertising looked like — a concentrated but broad ecosystem of 191 advertisers across 16 sectors, dominated by Retail Trade and Information, with delivery concentrated in prompts about purchasable products, health, fitness, and cooking. The authors expect the format to evolve along the arc of search advertising, in which Google took twelve years (2000 to 2012) to launch Google Shopping and progressively narrowed the visual and conceptual gap between organic and sponsored results — a trajectory they expect to accelerate given that ChatGPT may exceed 1 billion active users.

Future Directions

  • Continued monitoring of the income gradient. The finding that lower-income accounts are more likely to receive ads echoes longstanding concerns about differential advertising burdens, and the authors call for tracking it as OpenAI's targeting infrastructure matures.
  • Re-testing the race null result at greater scale. With 8–12 accounts per demographic group and 16–19 exposed accounts per racial group, the study is underpowered; the absence of a significant effect does not rule out racial disparities at larger scale.
  • Testing whether content guardrails hold. The near-zero ad rates on health, mental health, and politics are consistent with OpenAI's stated exclusions, but the authors describe the distinctions as "somewhat tenuous and context-dependent" — prompts about building muscle surfaced diet and exercise ads, and work-life balance surfaced ads for smart watches and "AI work agents."
  • Assessing generalizability. The study covers one platform, one country, and a three-week window at the very start of rollout; OpenAI announced a pilot of ads in other countries in May 2026, and results are not expected to generalize across markets or later periods.

Target Audience

Platform accountability researchers and AI auditors (who will find the methodology and its failure modes directly useful), advertising and consumer-protection regulators, civil-society and journalism organizations tracking commercial content in LLM interfaces, and HCI and internet-policy scholars studying how advertising formats evolve. Advertisers and product teams inside LLM platforms may also benefit from the descriptive picture of the early advertiser ecosystem and the prompt categories that drive delivery. Some familiarity with regression analysis or audit-study methods helps, but the plain-language framing of the findings makes the paper accessible to policy readers as well.

Authors’ abstract

This paper presents the first empirical study of advertising content being rolled out in the user-facing online interfaces of large language models (LLMs). We systematically examine possible demographic differences in ad content shown to U.S. users of ChatGPT using a sock puppet audit methodology. We create and deploy 91 sock puppets in a 3x3 factorial design, using geolocation cues (account IP proxies and location-signaling prompts) to signal three racial/ethnic groups (Black, Hispanic, and White) and three income terciles (low, medium, and high). We conduct data collection starting in February 2026, collecting over 3,000 advertisements from 186 unique advertisers in response to 335 prompts on a range of realistic user queries. We find that accounts begin receiving ads 14 days after account creation, and that lower-income accounts, regardless of race, are more likely to receive ads. In this first phase of ChatGPT ads, the ads themselves skewed heavily towards consumer goods, directed users to a specific advertiser rather than a particular product, and were clearly separated from the LLM's response text, observations we anticipate will change as ads continue being integrated into LLM chat interfaces. We release a public, searchable archive of all collected advertisements. Finally, we discuss the implications of our findings, and conclude with methodological and theoretical recommendations for future empirical studies of LLM advertisements.

Read the original paper