Skip to content
AI.info

Research

Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning

Overview Research area: AI safety and ethics — specifically how large language models moderate sensitive-but-legal cultural content (challenged books), bridging AI safety research, information science

arXiv
2608.11806
Published
2026-08-12
Authors
Xucheng Yu, Emily Knox, Haohan Wang

AI summary

Overview

Research area: AI safety and ethics — specifically how large language models moderate sensitive-but-legal cultural content (challenged books), bridging AI safety research, information sciences, and library/intellectual-freedom scholarship.

Technical level: Intermediate. The paper is a large-scale empirical measurement study rather than a methods paper; it uses simple two-proportion z-tests and one formal exaggeration parameter (γ), but the corpus construction, prompt tiers, and annotation design require some familiarity with LLM evaluation practice.

Scope: A systematic comparison of how six frontier LLMs respond to 400 restricted versus unrestricted books across 17 prompt designs and 40,800 query–response pairs, showing that refusal has been replaced by warning language and hesitation as the operative moderation mechanism.

Note: the supplied paper content ends mid-sentence in the Limitations section, so the full limitations discussion is not available here.

What This Paper Is About

The paper asks how modern large language models handle sensitive-but-legal topics — using books that have been formally challenged in US schools and libraries as a controlled testbed — rather than whether they handle them at all. Prior work on LLM moderation has assumed refusal is the main tool and has focused on jailbreaking or bypassing that refusal. This study tests that assumption at scale and finds it largely obsolete for this content class: models almost never refuse, and instead differentiate their answers through cautionary framing.

Key Contributions

  1. A zero-refusal empirical result. Across 40,800 query–response pairs, only 30 responses (0.07%) were outright refusals, with no operationally meaningful difference between restricted and unrestricted books (Δ = 0.12 pp, p < 0.001; restricted 0.13% vs. unrestricted 0.01%). The authors argue this effectively invalidates the premise of jailbreaking research for this content class.

  2. Identification of warning language and hesitation as the operative moderation mechanisms. Warning language appears 8–15 percentage points more often for restricted books (p < 0.001) and hesitation markers 2–5 pp more often, replicated across all six models.

  3. A prompt-design sweep showing framing drives the size of differentiation. Of 17 prompts, the scenario-based, institutionally framed ban_context prompt yields the largest warning gap (+19.0 pp), while the direct explicit_content prompt inverts the effect (−1.5 pp).

  4. A three-cluster model taxonomy and a three-layer moderation framework. The six models split into direct warning (Claude, Gemini, Grok), conditional hedging (GPT-4o), and moderate warning (DeepSeek, Qwen), and the authors describe a hierarchy from refusal (<0.1%) through warning (+8–15 pp) to hesitation (+2–5 pp).

Main Findings

  • Near-zero refusal across all models. Combined refusal rate is 0.07% across 40,800 tests. Per-model restricted-book refusal rates: Claude Sonnet 4.5 0.15%, Gemini 2.5 Flash 0.12%, DeepSeek-V3 0.29% (the highest), Qwen-Plus 0.21%, GPT-4o 0.03%, and Grok-4.1-Fast 0.00%. Grok-4.1-Fast produced zero refusals for restricted books. The pattern holds for both Western and Chinese providers.

  • Warning language is the primary differentiation signal. Averaged over all 17 prompts and 400 books, warning rates for restricted books are 44.6% versus 33.3% for unrestricted books, an average gap of +11.2 pp. Gaps range from +8.4 pp (Claude Sonnet 4.5) to +15.3 pp (Gemini 2.5 Flash). Hesitation gaps average +3.2 pp and range from +1.9 pp (DeepSeek-V3) to +4.8 pp (GPT-4o).

  • Differentiation is in framing, not length. Response length differences across book types range from −162 to +227 characters against average response lengths of 2,500–5,000 characters, a negligible fraction.

  • Sexual content is the strongest content-level discriminator. It appears in 87.5% of restricted-book responses versus 44.5% for unrestricted books, a +43.0 pp gap (the abstract reports the range as +33–52 pp). LGBTQ+ themes are the sharpest binary marker: 24.6% versus 1.0%, a +23.6 pp gap, appearing roughly 25 times more frequently. Race/racism shows +16.4 pp, drugs/alcohol +5.5 pp, and religion +1.4 pp (not significant).

  • Violence runs in the opposite direction. Violence is mentioned more often for unrestricted books (76.0% restricted vs. 86.0% unrestricted, −10.0 pp), and strong language likewise (49.1% vs. 54.9%, −5.7 pp). The authors attribute this to canonical works such as War and Peace, Les Misérables, and A Tale of Two Cities depicting historical violence framed as instructive, and note it is consistent with the fact that violence is rarely a primary ground for challenging books in the United States.

  • LLMs exaggerate the restricted-to-sexual-content association. Using the representativeness heuristic, the paper computes R = 1.97 for LLM mention rates and exaggeration parameters of γ ≈ 0.22 (Claude) and γ ≈ 0.53 (GPT-4o), against assumed true prevalences of 0.65 (restricted) and 0.05 (unrestricted). Predicted mention rates are approximately 78% (Claude) and 97% (GPT-4o).

  • Prompt design determines measurable differentiation. Top-ranked prompts: ban_context (+19.0% warning gap, +6.5% hesitation gap), lesson_plan_req. (+18.8%), gatekeeping_test (+16.8%), underage_reader (+14.5%), parental_vetting (+13.5%, +9.8% hesitation). Bottom-ranked: explicit_content (−1.5%), scene_or_moment (+1.8%), content_summary (+2.7%). Prompt tiers were Baseline (4), Enhanced (7), and Pressure (6).

  • Three model strategy clusters. Claude Sonnet 4.5 (37.7% warning, 11.2% hesitation, +8.4 pp gap) as direct warning; GPT-4o (46.3%, 19.8%, +10.0 pp) as conditional hedging; Gemini 2.5 Flash (47.9%, 11.0%, +15.3 pp) and Grok-4.1-Fast (51.9%, 10.3%, +12.8 pp) as high-warning direct; DeepSeek-V3 (41.5%, 12.8%, +11.4 pp) and Qwen-Plus (42.0%, 11.4%, +9.4 pp) as moderate warning.

  • Social framing drives caution. Prompts embedding a concrete social actor produce warning gaps two to four times larger than abstract queries, and first-person ownership ("my child") elicits more differentiated warning language than third-person phrasing ("a concerned parent"). Directly querying for controversial content activates blanket hedging across all books, erasing discriminative signal.

Methodology in Plain English

Building a controlled book corpus. The authors paired 200 restricted books against 200 unrestricted books. The restricted set comes from American Library Association records spanning 2000–2023: the Annual Top Ten Most Challenged Books lists, the Top 100 Most Challenged Books by Decade compilations, and individual challenge records for 2020–2023. Each restricted title has at least one documented formal challenge in a US school or library jurisdiction. Recorded challenge reasons break down as sexual content 40%, LGBTQ+ themes 25%, race/racism 20%, politics 10%, religion 5%. The unrestricted set was drawn from literary canon and bestseller lists with no ALA challenge record, matched by publication era and genre.

Why "restricted" and not "banned." The authors are explicit that ALA records document challenges — formal requests to restrict or remove access — which do not always result in removal. They argue "banned" would overstate the legal status of many titles.

Prompt design. They wrote 17 prompts escalating in contextual specificity, from neutral factual queries to scenario-based requests embedding age, role, and parental or pedagogical responsibility. Each prompt was paired with every book in both sets: 200 × 17 = 3,400 pairs per model per book type, 3,400 × 2 = 6,800 pairs per model, for 40,800 total pairs across six models.

Models. Claude Sonnet 4.5 (Anthropic), GPT-4o (OpenAI), Gemini 2.5 Flash (Google), DeepSeek-V3 (deepseek-chat), Qwen-Plus (Alibaba Cloud), and Grok-4.1-Fast (grok-4-1-fast-non-reasoning, xAI), all accessed via public APIs at temperature = 0.7.

Annotation. Each response was labelled automatically via keyword matching and regular expressions for refusal (binary), warning language, hesitation markers, content mentions (sexual content, violence, profanity, drugs, race, religion, LGBTQ+ themes), and recommendation tone. Manual validation on a 10% random sample yielded 94% agreement with automatic labels (Cohen's κ = 0.88).

Analysis. All comparisons used two-proportion z-tests; differences described as significant clear p < 0.001 unless otherwise noted. The paper also formalises a representativeness score R(a|X⁺) = P(a,X⁺)/P(a,X⁻) and an exaggeration parameter γ, where γ = 0 reproduces the empirical group difference faithfully, γ > 0 exaggerates, and γ < 0 attenuates. Ground-truth content prevalence estimates came from ALA challenge records plus a manual content audit of a 20% sample of each book set.

Why This Matters

Impact on research. The 0.07% refusal rate undercuts a core premise of jailbreaking literature, which assumes a meaningful refusal barrier worth circumventing. The authors argue the more productive research question is not whether models will engage with sensitive-but-legal content, but how they frame that engagement. The finding that six models from six providers — including both Western and Chinese developers — share the same three-layer structure suggests a convergent design philosophy rather than provider-specific policy. The work also extends prior findings on norm inconsistency and framing-dependence of guardrails into a new content domain.

Real-world applications:

  • Educators and librarians vetting reading materials: the paper shows that how a question is framed ("My 14-year-old wants to read this — should I allow it?" versus an abstract query) substantially changes how cautiously the model answers, with parental_vetting producing a +13.5 pp warning gap.
  • Parents deciding on age-appropriate reading: first-person framing ("my child") amplifies caution relative to structurally identical third-person framing.
  • AI policy and compliance teams designing moderation guidelines, since warning-rate gaps of +8.4 to +15.3 pp are the operative lever rather than allow/refuse decisions.
  • Content governance researchers needing validated, externally annotated test sets: the ALA challenge record provides a documented, independently updated annotation of cultural controversy that does not depend on the researchers' own judgement of sensitivity.

Industry relevance. The results imply that product teams optimising for refusal reduction may be solving a problem that no longer exists for sensitive-but-legal content, while under-instrumenting the warning and hedging layers that actually shape user-facing behaviour. The provider-level differences — GPT-4o's high warning plus high hesitation profile (19.8% hesitation rate) versus the low-hesitation direct-warning profiles of Gemini and Grok — give a concrete vocabulary for competitive and regulatory comparison. The finding that GPT-4o over-applies the restricted-to-sexual-content association more aggressively than Claude (γ ≈ 0.53 vs. γ ≈ 0.22) also raises a fairness concern, since the model flags sexual content beyond what the underlying books support.

Future Directions

  1. Cross-cultural replication. The authors state explicitly that the corpus is US-context English-language ALA data, and that whether the same patterns hold for books restricted in China, India, or Germany is an open empirical question beyond the scope of this study.

  2. Testing whether exaggerated associations cause mislabelling. The paper notes that a model that has learned "restricted books contain sexual content" may flag sexual content even for restricted books where it is absent, and may do so more aggressively than the base rate justifies — a systematic bias described as having a predictable direction, but not directly measured at the individual-book level here.

  3. Prompt-design methodology for moderation research. The failure modes identified — explicit_content inverting the gap to −1.5 pp and controversy_justification producing uniform hedging — suggest a need for validated prompt instruments that elicit content-sensitive responses without triggering domain-blind safety behaviour.

  4. Explaining convergence across providers. The paper reports that Chinese-developed models sit at an intermediate level comparable to GPT-4o, a pattern it says cuts against the idea that content differentiation is driven by provider-specific regulatory context. What mechanism produces convergence across differently trained systems remains unanswered.

Target Audience

This paper is most useful to AI safety and alignment researchers studying content moderation beyond refusal, to AI ethics and policy scholars interested in how value judgements get embedded in model outputs, and to library and information science researchers and practitioners concerned with how AI systems mediate access to contested books. Educators, parents, and librarians who use LLMs to vet reading materials will find the prompt-framing results directly actionable. Industry practitioners working on guardrails, trust and safety, or model evaluation will benefit from the three-layer moderation framework and the provider-level strategy clusters. Readers need only moderate familiarity with LLM evaluation; the statistical methods are simple proportions and z-tests, though the corpus and prompt design require careful reading.

Authors’ abstract

As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We study this question through a large-scale, systematic experiment using restricted versus unrestricted books as a controlled testbed: 40,800 query-response pairs, 400 books, 17 prompt designs, and six frontier models spanning six AI providers (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-4.1-Fast). Our restricted set is drawn from the American Library Association's Most Challenged Books records (2000-2023); we use restricted rather than banned throughout because the ALA documents formal challenges-requests to remove or restrict access-which do not always result in outright bans. Our central finding is a zero-refusal phenomenon: modern LLMs decline to discuss restricted books in only 0.07% of cases, effectively invalidating the premise of jailbreaking research for this content class. Differentiation occurs instead through warning language (+8-15 percentage points, p &lt; 0.001) and hesitation markers (+2-5 pp), with sexual content mention rate as the strongest individual signal (+33-52 pp). We further identify systematic differences between providers and show that prompt framing alone shifts the warning-rate gap by up to 19 pp. These results indicate that LLM content policy has shifted from binary refusal toward calibrated, context-sensitive disclosure-a finding that holds consistently across Western and Chinese AI providers.

Read the original paper