Research
Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study
Overview Research area: Algorithmic auditing of social-media recommender and search systems, with multimodal large language models (MLLMs) used as automated content annotators. Domain: youth safety, h
- arXiv
- 2608.17583
- Published
- 2026-08-18
- Authors
- Hamidreza Saffari, Francesco Pierri
AI summary
Overview
Research area: Algorithmic auditing of social-media recommender and search systems, with multimodal large language models (MLLMs) used as automated content annotators. Domain: youth safety, harmful-content exposure, cross-national content moderation.
Technical level: Intermediate. The audit design (sockpuppet accounts, stratified sampling, Cohen's kappa validation) is methodological rather than architecturally novel; the MLLM component uses off-the-shelf APIs in three input configurations.
Scope: A two-stage audit of TikTok in France, Italy, and Sweden using twelve (country, age) persona accounts to measure harmful-content exposure under passive For-You-page scrolling and under active harm-keyword search, with Gemini 2.5 Flash validated as a scalable annotator.
What This Paper Is About
Independent audits of harmful content on short-video platforms are hard to run at scale because human video annotation is costly and moderation judgments shift across languages. The authors ask whether a multimodal language model can substitute for native-speaker annotators well enough to measure exposure across countries and age groups. They then use the validated model to compare what TikTok's For-You feed serves, and what its search endpoint returns, to personas aged 13, 16, 19, and 40 in France, Italy, and Sweden.
Key Contributions
-
Empirical cross-national exposure findings on TikTok across three countries (France, Italy, Sweden) and four age personas (13, 16, 19, 40). The headline result is that harm-keyword search spikes harmful exposure 1.5–7.5 times over passive scrolling, that the spike is temporary, that the age differences observed in France and Sweden flatten under search, and that Italy has the highest passive harm rate at every age.
-
An automated harm-annotation pipeline using Gemini 2.5 Flash as the annotator, validated against native-speaker labels on a 300-video reference subset and then applied to a 10% sample of the full 36,971-video corpus for approximately $50 in API spend.
-
Released artifacts: a 13-category harm taxonomy aligned with TikTok's Community Guidelines, the annotation schema and prompts, per-country keyword lists, and metadata for the 36,971-video corpus.
-
A model-selection study (Stage-1) comparing four MLLMs across three input conditions on the same 300-video reference set, reading aggregate agreement and per-country disagreement side by side.
Main Findings
-
No configuration reaches strong agreement with human annotators. Across the ten Stage-1 (model, condition) combinations, Gemini 2.5 Flash with eight sampled frames plus text (E3) is the strongest, but tops out at aggregate Cohen's κ = 0.42, and no (model, condition) clears moderate agreement in every country.
-
Visual input helps, and frames beat native video on cost. Within Gemini, agreement improves from text-only (E1, κ = 0.17) to native video (E2, κ = 0.38) to eight frames (E3, κ = 0.42). E3 outperforms native-video upload at approximately 2 times lower per-call cost. All non-Gemini configurations sit below κ = 0.30 in aggregate; GPT-4o-mini's eight-frame run is the strongest non-Gemini combination at κ = 0.29 and reaches moderate agreement on Italian content, so the ranking re-orders by country.
-
Search is the dominant exposure pathway. The within-account scroll-pre to SEARCH to scroll-post cycle shows the search endpoint returning 35–56% harmful content, a 1.5–7.5 times increase over scroll-pre in ten of twelve country–age combinations, with no visible search-time block or warning.
-
The search spike does not persist. Across all twelve combinations, scroll-pre and scroll-post confidence intervals overlap and the absolute change is bounded by a few percentage points, meaning the elevation is confined to the search session rather than drifting the recommendation feed.
-
The age gradient collapses under search. At the SEARCH endpoint, harm rates fall into a 35–56% band across all twelve combinations; the youngest French and Swedish personas reach 37–42%, within a few points of adult combinations in the same countries.
-
The largest lifts appear where there is headroom. Older French and Italian personas sit below 25% at scroll-pre and see 50–56% at SEARCH. The exceptions are IT-13 and IT-16, where scroll-pre is already at 34–44% harm, leaving no room to lift. Same-age French and Swedish combinations sit at 9–19% pre-probe.
-
Italy leads under passive scrolling at every age. Italy's cross-country gap widens from a 4-point spread at age 13 (FR 24.5%, IT 27.9%, SW 23.7%) to over 25 points at age 19 (FR 22.5%, IT 48.6%, SW 38.4%); Italian age-19 is the most-exposed case at 48.6%. France stays flat at about 23% across all four ages, while Sweden rises from 24% at age 13 to 38% at age 19.
-
Across all phases, Italy carries the highest estimated rate at all four ages, with IT-16 the most-exposed case in the audit at 40.4%. France ranges 23.6% to 32.4%; Sweden ranges 24.2% to 31.0%. A per-country precision/recall correction leaves this ordering intact at ages 13, 16, and 19, but at age 40 the corrected French rate (approximately 49%) overtakes the corrected Italian (approximately 38%) and Swedish (approximately 33%) rates.
-
Sexually suggestive content dominates flagged items in every (country, age) combination, from 21.7% of harmful items (SW-13) to 50.0% (IT-16), with the caveat that Stage-1 strict primary-subcategory agreement is only 29% (54% under primary-or-secondary matching).
-
E2 and E3 disagree systematically. E2 reports harm rates 3–12 percentage points higher than E3 (e.g. FR-16: 35.1% vs 28.1%; SW-13: 31.3% vs 24.2%). The two modalities agree on the binary verdict at κ = 0.614 over n = 5,108 paired items, with per-country κ between 0.60 and 0.63. E2 over-flags visual-cue categories (Sexually Suggestive, Shocking and Graphic, Nudity) while E3 picks up more dialogic Harassment and Bullying.
-
The policy-tier split complicates the "age-blind" reading. Under passive scrolling, the 18+-restricted tier rises with persona age in Italy and Sweden (IT 17.2 pp at age 13 vs 26.7 pp at 40; SW 8.6 vs 24.5 pp), the direction age gating predicts, while the universally prohibited tier does not fall for minors (SW-13 carries the audit's highest passive prohibited-tier rate at 15.1 pp). Under SEARCH, the age separation on the restricted tier disappears, with 13-year-old personas receiving 24–35 pp of 18+-restricted content.
-
Provider safety filters under-count the most explicit harms. Gemini refuses to score about 1.1% of Stage-2 inputs (E2: 57 of approximately 5,300, ≈ 1.08%; E3: 64 of approximately 5,300, ≈ 1.21%; Fisher exact p ≈ 0.55). Only 23 of the 64 E3 refusals are surfaced through a SEARCH keyword. Among those, the block rate is dominated by Nudity and Body Exposure (≈ 4.8%), followed by Sexually Suggestive Content (≈ 2.6%), Shocking and Graphic Content (≈ 0.7%), and Disordered Eating and Body Image (≈ 0.5%). A worst-case sensitivity bound shifts headline rates by at most 1–3 pp and preserves the country ordering and the IT-13/IT-16 ceiling effect.
-
The Italian pattern survives three narrowings. Sexually Suggestive content contributes 23.8 pp of Italy's 39.0% passive rate against 12.2 pp in France and 15.8 pp in Sweden. On passive items that also appear in another country's corpus, Italy's rate is 19.2%, comparable to France (26.4%) and Sweden (28.3%), while Italy-exclusive content sits at 41.7%. English-language videos served to Italian accounts are flagged at 32.4%, above France (19.5%) with Sweden in between (27.5%).
-
Account-clustered uncertainty does not change the picture. A hierarchical bootstrap that resamples accounts before videos leaves intervals essentially unchanged (median width ratio 1.00, at most 1.85 times on the smallest scroll-post cells). Clustered intervals separate SEARCH from scroll-pre in ten of twelve combinations and preserve the scroll-pre/scroll-post overlap in all twelve.
Methodology in Plain English
The team ran a two-stage audit. In Stage-1, they built a 300-video reference set drawn at random with 25 videos per (country, age) cell, labeled independently by two native-speaker annotators per country who then reconciled disagreements through joint resolution. Four MLLMs — Gemini 2.5 Flash, Qwen3-VL-32B, GPT-4o-mini, and Mistral Large 3 — were scored against those labels under three input conditions: text only (caption plus audio transcript, E1), native video upload (E2), and eight uniformly sampled frames plus text (E3). The winning configuration was carried forward.
In Stage-2, the winning auditor was applied to a phase-stratified 10% sample of the full corpus. Data came from persona accounts on two axes: three countries and four ages, giving twelve (country, age) combinations, each instantiated by three independent accounts. Accounts reported their age at registration, used the system locale for their country, and were routed through a VPN endpoint in the target country because TikTok keys recommendations to IP country. A Tampermonkey userscript captured API responses and emulated scrolling with randomized scroll amounts and inter-scroll delays drawn uniformly from 500–1000 ms.
Collection had two phases. The passive phase captured pure For-You-page scrolling between 30 December 2025 and 11 January 2026 over five consecutive days per country, yielding 14,093 unique videos. The active phase ran five days per country between 13 and 19 April 2026 and followed a scroll-pre, then SEARCH with native-language harm keywords, then scroll-post cycle, yielding 47,674 capture events covering 22,878 unique videos (FR 7,130; IT 7,531; SW 8,217). Altogether the phases cover 36,971 unique videos. Because the two phases were collected in separate windows, they are analyzed separately.
Harm was coded against a 13-category taxonomy aligned with TikTok's Community Guidelines, injected into the system prompt with one-sentence definitions. Seven categories were probe-able by short keyword — three native-language keywords each per country, 21 keywords total — fixed before any active data collection began. Annotators gave a three-way verdict (Harmful, Not Harmful, Video Not Available), a required primary and optional secondary category, and a required confidence rating (Low, Medium, High). The Stage-2 sample contained 1,817 passive videos and 4,251 active-session videos; after removing unavailable videos, refusals, and parse failures, there were 5,227 usable E3 verdicts, 5,248 usable E2 verdicts, and 5,108 paired videos.
Why This Matters
Impact on research. The paper tests whether MLLM-as-annotator is viable for platform audits and finds the answer is qualified: κ = 0.42 is cheap enough for full-corpus scale (~$50) but below the bar for a drop-in replacement on policy-edge items. It also supplies a reproducible protocol — persona design, stratified sampling, modality comparison, clustered uncertainty — for cross-national audits under regulatory frameworks such as the EU Digital Services Act, and a concrete demonstration that provider safety layers introduce category-correlated measurement bias.
Real-world applications:
- Regulators and DSA-compliance reviewers assessing whether per-language moderation capacity matches exposure levels in a given market.
- Platform trust-and-safety teams deciding where to invest: the audit points at search-time moderation rather than recommender-side amplification as the main leverage point.
- Public-health and youth-safety researchers who need a low-cost pipeline for repeated cross-national measurement.
- Civil-society auditors who can reuse the released taxonomy, prompts, keyword lists, and annotation schema.
Industry relevance. The cost profile is the practical headline: roughly $50 in API spend to audit a 36,971-video corpus, with annotation throughput rather than compute cost as the remaining scaling bottleneck. The paper also documents a concrete failure mode of automated moderation infrastructure — Gemini's refusals cluster on Nudity and Body Exposure (≈ 4.8%) and Sexually Suggestive Content (≈ 2.6%), exactly the categories the audit is designed to surface — which means vendors offering moderation APIs are systematically under-reporting the most explicit harms.
Future Directions
-
Run a Stage-2 annotation pass. The design cannot currently say whether E2 (native video) or E3 (eight frames plus text) is closer to human truth on the Stage-2 population, nor whether the Stage-1 κ = 0.42 transfers to the full Stage-2 distribution.
-
Discriminate the mechanisms behind the flat SEARCH age gradient. The authors propose an item-overlap analysis within (country, keyword) to separate age-insensitive platform retrieval from a keyword set whose returned pool simply overlaps across personas.
-
Identify the source of the Italian pattern. Whether Italy's higher rate reflects content supply or per-language moderation capacity is, in the authors' words, not identifiable from the outside. The moderator-allocation figures they cite (French approximately 620, Italian approximately 396, Swedish approximately 98, averaged across 2023–2024 reporting periods) suggest a testable link.
-
Test temporal and persona generalizability. Collection spans a single time-bounded window and programmatically controlled sockpuppet accounts that lack real-user behavioral richness; extending the window and varying persona behavior would test whether the patterns hold.
Target Audience
Researchers in computational social science and algorithmic auditing; platform trust-and-safety and content-moderation practitioners; regulators and policy analysts working on the EU Digital Services Act and youth-safety rules; NLP researchers interested in MLLM-as-annotator pipelines, multimodal moderation evaluation, and cross-lingual annotation reliability. The paper is most useful to readers who want a reusable audit protocol and cost estimate more than a new model architecture.
Authors’ abstract
Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas (13, 16, 19, 40), collecting 36,971 videos from passive For-You-page scrolling and active sessions that scroll, search for harm keywords, and scroll again. To scale annotation, we validate four multimodal LLMs against native-speaker labels on a 300-video reference set. Gemini 2.5 Flash with eight sampled frames plus text performs best (aggregate kappa = 0.42), at half the per-call cost of native-video upload, and we apply it to a 10% sample for approximately \$50 in total API spend across both modalities. Keyword search returns 35-56% harmful content, a 1.5-7.5x increase over the scrolling baseline in ten of twelve country-age combinations; the spike is temporary and flattens the age differences observed in France and Sweden. Under passive scrolling, Italy has the highest harm rate at every age, with Italian age-19 reaching 48.6%. Overall, MLLM-based auditing offers a scalable approach for cross-national youth-safety audits, while provider safety filters (1.1% refusal rate) under-count the most explicit harms.