Research
AI use in American newspapers is widespread, uneven, and rarely disclosed
AI use in American newspapers is widespread, uneven, and rarely disclosed Overview Research area: Computational linguistics / AI-generated text detection applied to journalism, with a focus on media t

- arXiv
- 2510.18774
- Published
- 2025-10-21
- Authors
- Jenna Russell, Marzena Karpinska, Destiny Akinode, Katherine Thai, Bradley Emi, Max Spero, Mohit Iyyer
AI summary
AI use in American newspapers is widespread, uneven, and rarely disclosedOverview
Research area: Computational linguistics / AI-generated text detection applied to journalism, with a focus on media transparency and factuality.
Technical level: Intermediate. The methods (commercial AI detector API, LLM-based topic classification, chi-squared and Fisher's exact tests) are accessible to readers with basic quantitative literacy, though the paper leans on domain knowledge about newsroom practices and AI detection.
Scope: A large-scale audit of 186,507 articles from 1,528 American newspapers and 44,803 opinion pieces from three national outlets, using the Pangram AI detector to measure how much published newspaper text is flagged as partially or fully AI-generated, where that use concentrates, and how rarely it is disclosed.
What This Paper Is About
Newsrooms have adopted generative AI quickly, but nobody had systematically measured how much AI-written or AI-assisted text actually reaches readers in published American newspapers. The authors fill that gap by running a high-precision commercial AI detector over a large, freshly collected corpus of newspaper articles and opinion pieces, then manually auditing a sample for disclosure and factual accuracy. The goal is to quantify the scale, distribution, and transparency of AI use in journalism rather than to accuse individual journalists or outlets.
Key Contributions
-
The first large-scale audit of AI use in U.S. newspapers. The authors assemble two datasets —
recent_news(186,507 articles from 1,528 newspapers, June 15 to September 15, 2025) andopinions(44,803 op-eds from the New York Times, Washington Post, and Wall Street Journal, August 2022 to September 2025) — totaling 251,442 articles, all labeled with the Pangram detector. -
A multi-dimensional breakdown of where AI use concentrates. The paper reports AI use by newspaper circulation, geography, ownership group, topic, and language, plus a separate analysis of how AI use in opinion pieces has grown over time (0.1% in 2022 to 3.4% in 2025, a 25x rise).
-
A manual disclosure and factuality audit. The authors hand-inspect 200 AI-flagged articles for disclosure language and separately review 100 AI-Generated and 100 Human-written articles for hallucinated claims, and they catalog AI policies across the 200 sampled outlets.
-
Released artifacts for follow-up work. The paper releases links to the articles studied (not full texts), analysis code, an interactive dashboard at ainewsaudit.github.io, and a commitment to periodically update the dashboard with new articles and disclosure annotations.
Main Findings
-
Roughly 9% of recent news articles are flagged as AI-involved. In
recent_news, 9.1% of articles are labeled by Pangram as either AI-Generated (5.2%) or Mixed (3.9%), with the remaining 90.9% classified as Human-written. -
AI use is higher at smaller, local outlets. Only 1.7% of articles at papers with circulation over 100K are flagged, versus 9.3% at papers below 100K (article level: χ²(1) = 1175.6, p < 10⁻²⁵⁰; newspaper level: 8.5% for smaller outlets vs. 5.0% for very large outlets, Welch's t(≈23) = 2.24, p = 0.032, d = 0.22). The authors suggest large national papers enforce stricter editorial constraints on automation.
-
AI use varies geographically. Rates peak in the mid-Atlantic and southern U.S. — Maryland (16.5%), Alabama (13.9%), Tennessee (13.6%) — and are lowest in the Northeast, including New Hampshire (2.9%) and Massachusetts (3.4%).
-
Topic matters. Weather articles show the highest average AI likelihood (27.7%), followed by science and technology (16.1%) and health (11.7%). Lower rates appear for conflict and war (4.3%), crime, law and justice (5.2%), and religion (5.3%).
-
Ownership groups differ sharply. Boone News Media and CherryRoad Media both exceed 15% detected AI use, while Nash Holdings and Hearst Corporation fall under 1%. Patterns are topic-specific: Advance Publications shows 81% AI use for weather reporting, Boone Media reaches 67.7% in science and technology and 55.7% in lifestyle and leisure, CherryRoad Media 37.8% in sports, and Lee Enterprises 15.4% in foreign policy coverage.
-
Non-English articles show far more AI use. Only 8.0% of English-language articles are flagged as AI-Generated or Mixed, versus 31.0% for articles in other languages. About 80% of non-English AI use comes from U.S.-based Spanish-language reporting (7.2K articles); other major languages include Portuguese (468), Vietnamese (403), French (343), and Polish (314).
-
Machine translation is not the main driver of the non-English gap. Translating 300 English articles into 12 languages with GPT-4.1 lowered AI-likelihood scores by an average of 13%; binary agreement with the English originals was 83.2%, and translation was far more likely to suppress AI predictions than introduce them (McNemar's exact test, FDR-corrected p < 0.001 for all but Japanese). Vietnamese showed the largest decrease (Δ = −0.21) and lowest agreement (74.7%), while Portuguese and Spanish retained ≥88% agreement.
-
Most AI involvement is partial, not full replacement. Of the 17,059 articles detected as using AI, 42.7% are labeled Mixed and 57.3% AI-Generated. At the author level, 1,453 of 34,608 writers produced at least some AI content; 54.8% of those primarily publish mixed articles, while 36.1% rely mostly on AI-generated text.
-
Disclosure is rare. In a sample of 200 AI-flagged articles from unique newspapers, 96.5% of authors and 94.0% of publishers did not disclose AI use (95% Wilson CIs: [93.0%, 98.3%] and [89.8%, 96.5%]). Only 7 of the 200 articles disclosed AI use. Of the 200 outlets examined, 12 have policies allowing AI, 2 prohibit it, and 186 have no public policy.
-
AI-flagged articles are far more likely to contain hallucinations. Manual review of 100 AI-Generated and 100 Human-written articles found that 41% of AI-labeled news contained at least one hallucinated claim, versus 5% of human-written articles — an 8.2x difference (Fisher's exact test, p = 2.3 × 10⁻⁹). Typical errors include fabricated quotes, incorrect statistics, and misdated events.
-
Opinion pieces at top papers carry more AI use than news. Across NYT, WaPo, and WSJ during June–September 2025, AI use appears in 4.56% of opinion pieces versus 0.71% of other articles from the same outlets — 6.4 times more likely (n = 3,420 opinions; n = 10,129 all articles). By outlet: Washington Post 5.51% vs. 0.55%, WSJ 4.99% vs. 0.74%, NYT 2.94% vs. 1.80%. Mixed authorship dominates in both settings (86.5% of AI use in opinions, 86.1% in non-opinion articles).
-
AI use in opinion pieces has risen roughly 25x since 2022. The share flagged rises from 0.1% in 2022 to 3.4% in 2025. By outlet: WSJ 0.1% to 3.4%, Washington Post 0.2% to 4.3%, NYT 0.0% to 2.6%. As a pre-ChatGPT sanity check, only 5 of 5,029 opinion articles published before December 2022 were flagged, implying an empirical false-positive rate of 0.10% (95% Wilson CI: 0.04%–0.23%).
-
Guest contributors drive opinion-page AI use. Of the 219 unique authors in
opinionswith at least one AI-flagged article, most are infrequent contributors rather than full-time journalists. Political figures, executives, and scientists show the highest AI use, with many using AI for all their articles, while veteran opinion columnists show near-zero incidence (<0.5%). -
Reporters tracked over time show steep AI adoption. In the separate
ai_reportersdataset, reporters increase their AI use from approximately 0% prior to 2023 to over 40% in 2025 on average. -
AI use in opinion writing has spread across topics. Politics and government dominate opinion content (56.9%, versus 15.0% in
recent_news). Year-over-year growth in AI use is largest in crime and law (about 31x), religion (16x), and economy and business (14x), while science and technology leads in absolute terms (1.2% to 9.0%, roughly 9x). Gains also appear in conflict, war and peace (12x), human interest (10x), and politics and government (9x).
Methodology in Plain English
The authors built two corpora. For recent_news, they started from a list of 6,175 U.S. newspaper URLs, filtered down to 1,528 reachable sites with active RSS feeds, and roughly twice a week from June 15 to September 15, 2025 automatically downloaded up to 50 recent articles per paper, cleaning the text with Trafilatura and Newspaper4K. For opinions, they pulled full text and metadata for op-eds from the New York Times, Wall Street Journal, and Washington Post via ProQuest Recent Newspapers — 16,964 WSJ, 15,977 WP, and 11,862 NYT articles.
Every one of the 251,442 articles was passed through Pangram (v2.0), a commercial AI detector that returns both a 0–100% AI-likelihood score and a categorical label. The authors collapse Pangram's fine-grained labels into three buckets: Human-written, Mixed, and AI-Generated. Pangram reports a false-positive rate of about 0.001% on news text, and an independent study (Jabarian and Imas, 2025) measured 0.08%. As a cross-check, a second commercial detector (GPTZero) agreed with Pangram on 88.2% of a balanced 1,000-article held-out set (Cohen's κ = 0.764), with 98.4% agreement on the human subset and 78.4% on the AI subset.
Each article was also assigned one of 19 topics (the 17 top-level IPTC Media Topics categories plus "Other" and "Obituary") using zero-shot prompting with Qwen3-8B. Two authors independently re-labeled 100 sampled articles, achieving 87% inter-annotator agreement (κ = 0.85) and 77% agreement between the model and human labels. Newspapers were linked to print circulation data from the U.S. News Deserts Database (matching 54.6% of articles and 49.7% of newspapers) and to ownership data from the Medill Local News Initiative (covering 87% of recent_news).
For disclosure, the authors manually inspected 200 AI-flagged articles from unique newspapers and checked 200 outlet websites for public AI policies. For factuality, they manually compared 100 AI-Generated and 100 Human-written articles for hallucinated claims. Statistical comparisons use chi-squared tests, Welch's t-tests, Fisher's exact test, McNemar's exact test, and Wilson confidence intervals.
Why This Matters
Impact on research. This is the first audit of AI use at this scale in U.S. newspapers, establishing an empirical baseline (roughly 9% of recent articles flagged) that future studies can compare against. It also contributes a validated methodology — detector plus topic classifier plus manual disclosure and factuality audits — that can be ported to other media domains, and it surfaces a striking asymmetry: AI use is most common in short-form, high-volume content such as weather, science and technology, and health, and in opinion pieces by guest contributors rather than staff journalists.
Real-world applications:
- Editorial policy design: Newsrooms can use the mixed-authorship framework the authors propose — AI uses acceptable without disclosure (grammar, style), acceptable with disclosure (summarization, rewrites), and not permitted (full article generation).
- Contributor vetting: Because opinion pieces by guest contributors show the highest AI incidence, outlets can require author attestations about AI use at submission and screen submissions for AI cues.
- Reader-facing transparency: Platforms and publishers can adopt labeling standards that let readers distinguish AI-assisted editing from AI-generated content.
- Detection tooling and benchmarking: The released dashboard, code, and dataset links give detection researchers real-world, labeled news text to test multilingual and domain-specific false-positive rates.
Industry relevance. The findings land directly on a live industry debate. The authors note the irony that several media groups suing AI companies over training on their content — citing Advance Local Media v. Cohere — are simultaneously heavy users of LLM-generated articles. Pew Research figures cited in the paper frame the stakes: 49% of Americans who get news directly from AI assistants report encountering inaccurate information, 56% would feel less confident about a news article if they knew an AI wrote it, and 76% believe it is extremely important to know whether text they read is AI-generated. Prior survey work cited (Newman et al., 2025) found comfort with AI-generated material is low (19% for AI-Generated, 30% for Mixed), and a BBC/EBU study found around 55% of AI-generated responses had accuracy issues — consistent with the 41% hallucination rate reported here.
Future Directions
-
Moving beyond disclosure audits to intent. The authors cannot determine what role AI played in Mixed articles, so a key open question is how to separate editorial assistance (style edits, summarization) from content generation in a principled, auditable way.
-
Extending the factuality analysis. The paper only checks hallucination rates on 100 AI-Generated and 100 Human-written articles and does not perform a large-scale evaluation of the factuality of AI-Generated articles beyond quotation authenticity. Scaling this audit is the obvious next step.
-
Better multilingual detection. Because non-English content shows 31.0% flagged AI use versus 8.0% for English, and machine translation systematically suppresses AI signals, the authors treat their cross-lingual findings as exploratory. The field needs detector validation across more languages, human translation, and alternative translation models.
-
Tracking adoption over time and broadening geographic coverage. The authors commit to periodically updating their dashboard with new articles and disclosure annotations. Their dataset is U.S.-focused, RSS-dependent (excluding print-only papers), and weighted toward regional and local outlets, so whether these patterns hold for global journalism remains an open question.
Target Audience
This paper is most valuable to journalism and media studies researchers, newsroom editors and standards committees formulating AI policies, AI detection and text provenance researchers, and policy analysts concerned with transparency in automated content. Communication scholars studying opinion-page influence and readers interested in how AI is reshaping the news they consume will also find the disclosure and hallucination findings directly relevant. The statistical methods are accessible to an intermediate audience, but the paper explicitly cautions that all findings are detector outputs and should not be read as authorship attributions or judgments about specific journalists, outlets, or companies.
Authors’ abstract
AI is rapidly transforming journalism, but the extent of its use in published newspaper articles remains unclear. We address this gap by auditing a large-scale dataset of 186K articles from online editions of 1.5K American newspapers published in the summer of 2025. Using Pangram, a state-of-the-art AI detector, we discover that approximately 9% of newly-published articles are either partially or fully AI-generated. This AI use is unevenly distributed, appearing more frequently in smaller, local outlets, in specific topics such as weather and technology, and within certain ownership groups. We also analyze 45K opinion pieces from Washington Post, New York Times, and Wall Street Journal, finding that they are 6.4 times more likely to contain AI-generated content than news articles from the same publications, with many AI-flagged op-eds authored by prominent public figures. Despite this prevalence, we find that AI use is rarely disclosed: a manual audit of 100 AI-flagged articles found only five disclosures of AI use. Overall, our audit highlights the immediate need for greater transparency and updated editorial standards regarding the use of AI in journalism to maintain public trust.