Research
Beyond the Rabbit Hole: Mapping the Relational Harms of QAnon Radicalization
Overview Research area: Computational social science and natural language processing applied to online radicalization, specifically the study of QAnon's effects on the families and friends of believer

- arXiv
- 2601.17658
- Published
- 2026-01-25
- Authors
- Bich Ngoc Doan, Gianmarco De Francisci Morales, Giuseppe Russo
AI summary
Overview
Research area: Computational social science and natural language processing applied to online radicalization, specifically the study of QAnon's effects on the families and friends of believers.
Technical level: Intermediate. The paper combines accessible narrative analysis with moderately technical methods (topic modeling, mixed-membership clustering, compositional data transformation, logistic regression), but the framing and findings are written for a broad audience.
Scope: The paper analyzes 12,747 first-hand Reddit posts from r/QAnonCasualties to derive six recurring "radicalization personas" and show that these personas predict which specific negative emotions the people close to the believer report.
What This Paper Is About
Most computational research on conspiracy theories studies the believers themselves by tracking their online activity. This paper flips the perspective and studies the harm felt by the relatives and close acquaintances of QAnon believers, who write about the experience in the support community r/QAnonCasualties. The goal is to build a computational pipeline that identifies how radicalization unfolds, group those trajectories into distinct archetypes, and then quantify which emotional tolls (anger, fear, sadness, disgust) are associated with each archetype.
Key Contributions
-
QAnonCasualties-12k dataset. The authors release the first publicly available dataset built from 12,747 first-hand narratives documenting offline radicalization into QAnon, annotated with thematic traits describing the believer's trajectory and the narrator's emotional response. To protect privacy, they release Reddit post identifiers and derived annotations (thematic traits, persona mixtures, emotion labels, narrator control variables) rather than raw post text. Code and data are available at https://github.com/ngoccc/qanoncasualties-analysis.
-
Six data-driven radicalization personas. The authors derive the first empirical typology of offline radicalization journeys as witnessed and reported by close relations.
-
Evidence that personas predict specific emotional harms. Using regression on persona composition, the paper provides large-scale quantitative evidence that which archetype a believer fits predicts which emotion the narrator reports, reframing radicalization as a measurable relational phenomenon.
Main Findings
-
Conservative identity is the dominant precursor. Among pre-radicalization conditions, Conservative Political Identity appears in 1,489 profiles, nearly 50% more than the next most common factor, despite substantial heterogeneity in psychological, socio-economic, and demographic vulnerabilities.
-
Triggers cluster around the pandemic media environment. COVID-19 Lockdowns (2,402 profiles) dominate by a wide margin, followed by Religious Influence (1,181) and platform-specific media exposure (Fox News, Facebook, YouTube Algorithms). Political events such as Trump's Election play a comparatively minor role (181 profiles). A considerable share of profiles (1,714) show no clearly identifiable trigger, labeled "Unknown."
-
Post-radicalization traits split into two largely independent axes. A relational axis (Social Deterioration, Interpersonal Harm) plus Lifestyle and Media Use changes, and an ideological axis (Child Harm, Apocalyptic Narratives, Pro-Trump and War-Inducing Rhetoric, Medical and Science and Technology Mistrust). The two axes co-occur in only 32.6% of profiles, suggesting behavioral and epistemic radicalization are partially decoupled.
-
Six personas capture overlapping, not exclusive, pathways. Profiles occupy on average 1.70 personas (with a threshold of 1/k), and 52.2% show mixed membership. The six personas are: P1 The Health-Triggered Conspiracy Theorist, P2 The Political Extremist, P3 The Social Media Spiral, P4 The Religious Apocalypticist, P5 The Conservative Identity Protector, and P6 The Pandemic-Triggered Skeptic.
-
Christian symbolism and Trumpism fuse in over 1,000 believers, surfacing as co-loading on the Religious Apocalypticist and Political Extremist personas — an interaction effect invisible to single-mechanism studies.
-
Anger and disgust track radicalization perceived as an active, deliberate transgression. Profiles dominated by Dispositional personas (rooted in pre-existing political or religious beliefs) are more likely to evoke anger (OR = 0.93) and disgust (OR = 0.87) than those driven by Situational factors. Within the Situational group, the Chronic Personal Crisis persona P1 is linked to more anger (OR = 1.08) and disgust (OR = 1.09) than the Acute Public Crisis persona P6.
-
Disgust responds to specific moral content. The Outward Aggression persona P2 corresponds to stronger disgust (OR = 0.90) than the Inward Identity Defense persona P5, and the Digital-based Social Media Spiral P3 shows higher odds of disgust (OR = 0.88) than the Corporeal-based personas P1 and P6.
-
Fear and sadness track perceived cognitive and personal collapse. The Corporeal-based personas P1 and P6 have a stronger association with both fear (OR = 1.09) and sadness (OR = 1.11) than the Digital-based P3. Fear is more prominent in narratives of the Sacred Religious Apocalypticist P4 (OR = 0.90), while sadness is more prevalent in those of Secular personas P2 and P5 (OR = 1.11).
Methodology in Plain English
The authors collected 31,837 posts from r/QAnonCasualties, spanning the period from the subreddit's launch on July 3, 2019, through December 31, 2024. They removed automated accounts, duplicates, posts shorter than 50 words, and posts that did not describe an interpersonal dynamic — meaning they had to mention both a narrator (the post author, captured by terms like "I" or "me") and an object of narration (the believer, captured by terms like "mother," "husband," or "friend"). Three annotators reviewed a random sample of 200 retained posts, finding that 87% satisfied this dual-perspective requirement (Fleiss' kappa = 0.71). The final dataset contains 12,747 posts.
To identify what the posts were about, they applied BERTopic (using UMAP for dimensionality reduction and HDBSCAN for clustering, with min_cluster_size set to 50). This produced 461 topics. They benchmarked BERTopic against LDA and an LLM-based prompt topic model at k = 50 topics and found BERTopic had the highest diversity (0.700), coherence (-0.030), and stability (ARI = 0.873) versus LDA (0.317, -0.064, 0.135) and the LLM-based method (0.092, -0.136, 0.533). HDBSCAN flagged 128,667 outlier sentences (35.3%), which were discarded, leaving 235,694 sentences.
Because many of the 461 topics were conversational artifacts rather than descriptions of the believer, the authors built an "LLM Co-annotation Framework": they prompted gpt-4o-mini with few-shot exemplars to mark topics relevant if they concerned belief adoption, behavioral change, or interpersonal dynamics, with the first author annotating all topics and disagreements resolved with the co-authors. This yielded 198 relevant topics.
They then mapped those 198 topics onto the three-phase radicalization model of Borum (2011), operationalized by Klausen (2016): Pre-Radicalization Conditions, Trigger, and Post-Radicalization Characteristics. After merging conceptually similar topics (especially in the dense third phase), they obtained 50 thematic traits, and each post became a "profile" — the subset of traits it expresses.
To find personas, they used Latent Dirichlet Allocation in a mixed-membership formulation, so each profile is a distribution over personas rather than belonging to one. They swept k from 4 to 15 and selected k = 6 based on topic coherence and qualitative review.
For the second research question, they fit one binary logistic regression per emotion (anger, fear, sadness, disgust, drawn from Plutchik's framework). Emotions were labeled with gpt-4o-mini; validation on a 100-post sample gave mean kappa = 0.6 against two independent annotators, outperforming NRC and GoEmotions-finetuned baselines. They controlled for the narrator's relationship to the believer, gender, and age, extracted via regular expressions plus gpt-4o-mini zero-shot classification (F1 of 0.95 for relationships, 0.77 for age, and 0.70 for gender). Because persona probabilities are compositional (non-negative and summing to one), they applied an Isometric Log-Ratio transformation, producing five "balances" defined by a Sequential Binary Partition that contrasts personas thematically (for example, Situational versus Dispositional, or Corporeal versus Digital).
Why This Matters
Impact on research. The paper shifts the unit of analysis in radicalization research from the individual believer to the relational network around them, and provides a reusable modular pipeline — contextual topic modeling, graphical modeling, and LLM-assisted classification — for extracting narrative structure from unstructured text. It also gives the first large-scale quantitative test of how radicalization profiles map onto specific emotional harms, and provides a quantitative prevalence estimate for mechanisms that prior work had only identified in isolation.
Real-world applications.
- Designing psychosocial support programs tailored to the specific distress families report, since some pathways evoke anger and disgust while others evoke fear and sadness.
- Training support-group moderators and counselors to recognize which radicalization patterns predict which relational harms.
- Informing de-radicalization and family-intervention efforts by giving them a vocabulary of trajectories rather than a single "rabbit hole" narrative.
- Providing a starting feature set for models that might forecast outcomes such as relationship severance or receptiveness to de-radicalization.
Industry relevance. Social media platforms and moderation teams could use persona-level analyses to understand how algorithmic exposure (YouTube, Facebook, 4chan) maps to downstream relational harm. The LLM co-annotation and LLM-assisted emotion detection methods are also directly relevant to applied NLP teams building annotation pipelines where scale and cost matter.
Future Directions
- Build a richer emotional taxonomy. The current study relies on four pre-defined negative emotions and misses context-specific feelings such as helplessness, irony, or betrayal.
- Scale up human validation. The current approach trades scale for precision, relying on limited, largely disagreement-triggered human review; larger validation would also support custom classifiers better able to capture linguistic markers of trauma and support-seeking.
- Address self-selection bias. Because all data comes from r/QAnonCasualties, the observed distribution of pathways may not represent everyone affected. The authors propose extending to other platforms and media.
- Test generalization and temporal dynamics. Future work should analyze data longitudinally to model how trajectories evolve, and test whether the personas generalize to conspiracy movements beyond QAnon.
Target Audience
Researchers in computational social science, NLP practitioners working on narrative or emotion analysis, and scholars of radicalization and extremism will benefit most. The paper is also valuable for mental health professionals, family-support practitioners, and platform trust-and-safety teams who need an evidence-based framework for understanding how conspiracy belief harms the people around the believer rather than the believer alone.
Authors’ abstract
Large-scale computational research on conspiracy theories has focused exclusively on believers' online behavior, leaving the harm experienced by those closest to them under-examined. This paper bridges this gap by analyzing 12747 stories from r/QAnonCasualties, an online support group for people who have ``lost'' someone to conspiracy beliefs. We design a computational pipeline to extract fine-grained thematic traits from personal narratives and cluster them into six coherent radicalization personas, which we then link to the emotional toll reported by narrators via LLM-assisted emotion detection and regression modeling. We find that personas are meaningful predictors of specific emotional harms: radicalization perceived as a deliberate ideological choice is associated with anger and disgust, while personas marked by personal and cognitive collapse correspond to fear and sadness. This work provides an empirically grounded computational framework for understanding the relational harms of radicalization, opening new avenues for research into its wider social consequences.