Research
SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection
Overview Research area: Natural Language Processing, specifically abusive-language and cyberbullying (CB) detection, synthetic data generation with large language models, and dataset construction. Tec

- arXiv
- 2511.11599
- Published
- 2025-10-30
- Authors
- Arefeh Kazemi, Hamza Qadeer, Joachim Wagner, Hossein Hosseini, Sri Balaaji Natarajan Kalaivendan, Brian Davis
AI summary
Overview
Research area: Natural Language Processing, specifically abusive-language and cyberbullying (CB) detection, synthetic data generation with large language models, and dataset construction.
Technical level: Intermediate. Readers should be comfortable with concepts such as LLM prompting, inter-annotator agreement (Cohen's and Fleiss' kappa), F1 scores, and BERT-style classification, though the paper explains each metric it uses.
Scope: The paper introduces SynBullying, a publicly available multi-LLM synthetic conversational dataset for cyberbullying detection, and evaluates how faithfully its three constituent LLM-generated subsets reproduce an authentic teen role-play cyberbullying corpus, along with their usefulness for training and augmenting CB classifiers.
What This Paper Is About
Cyberbullying is inherently conversational: whether a message is harmful often depends on what came before it, who said it, and how the exchange escalated. Most existing CB and toxicity datasets are message-level, lacking this context, and collecting genuine conversational data from minors is constrained by ethical, legal, and psychological barriers. This paper builds a synthetic substitute: a multi-turn, role-based, fine-grained labelled conversational CB dataset generated by three different LLMs, then systematically checks how realistic and how useful that synthetic data actually is compared with authentic data.
Key Contributions
-
SynBullying dataset release. A synthetic, multi-turn conversational cyberbullying dataset generated by three LLMs (GPT-4o Feb-2025, Llama-3.3-70B-Instruct, and Grok-2 Feb-2025), each contributing 40 conversations, with every message automatically annotated by GPT-4o (Sept-2025 version). It is publicly available at https://huggingface.co/datasets/arrkaa-NLP/SynBullying.
-
Context-aware, fine-grained annotation. Rather than annotating isolated posts, labelling is performed over whole conversations so that harmfulness is judged within conversational flow, and harmful messages receive one or more of 12 fine-grained CB type labels aligned with the taxonomy of Van Hee et al. (2015) (the "Defense" category is excluded).
-
Reliability validation of an LLM annotator. GPT-4o's labels are compared against human gold labels on the authentic dataset for binary harmfulness and per-CB-type classification, including a direct comparison with human–human agreement.
-
Multi-dimensional comparative and utility analysis. Synthetic data is compared with authentic data on lexical/conversational statistics, sentiment and toxicity, role dynamics, harm intensity, and CB-type distribution, and is then tested as standalone training data, as a transfer target, and as augmentation for a BERT-based classifier.
Main Findings
-
GPT-4o approximates human-level binary annotation. Against human gold labels on the authentic dataset, GPT-4o achieved accuracy 0.833, precision 0.688, recall 0.825, and F1 0.750, with Cohen's kappa = 0.627 and Fleiss' kappa = 0.625 (substantial agreement). The human–human Cohen's kappa for this task is reported as 0.69. The model slightly over-predicts the harmful class and occasionally misclassifies borderline, sarcastic, or context-dependent instances.
-
Fine-grained CB type labelling varies sharply by category. GPT-4o's best F1 scores were on explicit categories: Insult_Attacking_Relatives (F1 = 71.43) and Insult_Body_Shame (F1 = 66.02). Moderate performance appeared for Threat_or_Blackmail (63.49) and Curse_or_Exclusion (60.96). The hardest were Encouragement_to_Harassment (47.54) and Defamation (11.11). For Defamation, human–human agreement is also kappa = 0.0, and human kappa values ranged from 0.00 for Defamation to 0.66 for Insult, indicating intrinsic annotation difficulty rather than a model deficiency. GPT-4o reached comparable or higher reliability than humans on Threat_or_Blackmail, Curse_or_Exclusion, and Encouragement_to_Harassment.
-
No single LLM matches authentic data on all dimensions. LLaMA most closely matched the authentic corpus on lexical diversity (MTLD 51.55 versus authentic 56.27; Grok 51.69; GPT-4o 75.54) and average tokens per message (15.79 versus 8.34 authentic; Grok 16.51; GPT-4o 9.30). GPT-4o had the largest vocabulary (3316) and highest MTLD, reflecting stylized, over-diverse output. GPT-4o additionally deviated most on role balance, equalizing bully and victim participation (bully 18.22%, victim 14.65% versus authentic bully 28.74%, victim 16.01%) and amplifying VictimSupport (43.40%).
-
Toxicity and profanity profiles diverge in opposite directions. Authentic data had 19.21% toxicity and 0.78 profanity per 100 tokens. GPT-4o was strongly sanitized (1.95% toxicity, 0.03 profanity, 48.72% positive sentiment). Grok exaggerated hostility (39.80% toxicity, 1.36 profanity, 45.88% negative sentiment). LLaMA sat in between (15.42% toxicity, 0.26 profanity). Sentiment was measured with NLTK's VADER analyzer; toxicity with ToxicBERT (messages scoring at or above 0.5 flagged as toxic).
-
Harmful/harmless balance: authentic 36.36% harmful; GPT-4o 34.53%; Grok 53.56%; LLaMA 38.64%. Jensen–Shannon divergence from authentic was 0.0002 for GPT-4o (p = 0.144), 0.0003 for LLaMA (p = 0.0937), and 0.0150 for Grok (p = 3.81e-38). GPT-4o and LLaMA therefore closely reproduce the authentic harmful/harmless balance, while Grok overestimates harm.
-
CB type distributions differ significantly across all models. GPT-4o was closest to authentic (JSD = 0.0317), followed by Grok (0.0739) and LLaMA (0.0803), but all three were significantly different from authentic data (p < 1e-28). Insult_General is the most frequent type in all datasets; Grok and LLaMA over-represent discrimination-related insults, while GPT-4o better preserves secondary types such as Curse_or_Exclusion and Defamation.
-
Synthetic-only training transfers poorly to authentic conversations. A BERT-base-uncased classifier (110M parameters, with a linear classification layer, repeated at least ten times per experiment) trained only on synthetic data reached 53.4 Harm-F1 for LLaMA, 50.5 for Grok, and 49.0 for GPT-4o when tested on authentic data, against an authentic-only baseline of 68.2 Harm-F1 and 80.7% accuracy.
-
Reverse transfer exposes blind spots. Authentic-trained models tested on synthetic content scored 75.7 Harm-F1 on Grok messages, 61.7 on LLaMA, and only 37.9 on GPT-4o, showing that authentic-only classifiers largely miss GPT-4o-style harmful content.
-
Augmentation restores near-baseline performance. Adding synthetic data to authentic training data yielded 68.7 Harm-F1 for Auth+LLaMA, 67.7 for Auth+Grok, and 65.8 for Auth+GPT when tested on authentic conversations, compared with the 68.2 baseline.
-
Role-specific utility recommendation. LLaMA is described as most effective for realistic augmentation, Grok for adversarial worst-case evaluation, and GPT-4o for exposing system blind spots.
Methodology in Plain English
The researchers started from an existing authentic cyberbullying corpus created through teen role-play sessions (Sprugnoli et al., 2018; English version by Verma et al., 2023), in which participants take the roles of Victim, Bully, Bully Supporter, and Victim Supporter, and conversations are triggered by one of four predefined scenarios labelled A–D. They benchmarked all synthetic data against this corpus.
For generation, they wrote a role-based prompt assigning eleven fictional teenage participants (one victim, two bullies, four victim supporters, and four bully supporters) and asked three LLMs to produce realistic, profanity-rich multi-turn conversations based on the same A–D scenarios. The task was framed as academic work for cyberbullying detection; when a model refused to produce harmful content, the prompt was reissued until the target number of conversations was obtained. The prompt template follows Kazemi et al. (2025), and prompts are documented in the paper's appendix.
For labelling, GPT-4o was chosen as the annotator after initial evaluations showed it performed best on CB-related labelling among the tested models. Crucially, each LLM call processed an entire conversation and returned labels for every message, so annotations could consider conversational context. Each message received an is_harmful yes/no label and, where applicable, one or more CB type labels. Before applying this at scale, GPT-4o's reliability was measured against human gold labels on the authentic dataset and compared with human–human agreement.
Comparative analysis covered five dimensions: lexical and conversational statistics (MTLD, vocabulary size, messages per conversation, tokens per message), sentiment and toxicity, role-based message distribution, harm intensity, and CB type distribution. Sentiment used VADER with standard thresholds (at or above 0.05 positive, at or below -0.05 negative); toxicity used ToxicBERT with a 0.5 threshold; profanity used a curated lexicon handling censored variants such as f*k. Distributional similarity was quantified with Jensen–Shannon divergence and p-values.
Finally, classifier experiments used BERT-base-uncased with a linear binary classification layer, run at least ten times with different random initializations and averaged, under leave-one-conversation-out cross-validation because the authentic set contains only 10 conversations. Four experiment groups were defined: authentic-only baseline, synthetic-only training, authentic-trained models tested on synthetic content, and augmentation combining both.
Why This Matters
Impact on research. The paper provides a scalable, ethically safer route to conversational cyberbullying data at a time when message-level datasets dominate and authentic conversational data from minors is hard to collect. It also contributes a methodological template: validate the LLM annotator against human gold labels before trusting it at scale, and report where LLM–human agreement is comparable to or exceeds human–human agreement. Its finding that purely synthetic training data cannot substitute for authentic annotations is a caution for the growing synthetic-data literature.
Real-world applications.
- Content moderation systems for social platforms and messaging apps, including detection of multi-turn escalation rather than single messages.
- Adversarial stress-testing of existing moderation pipelines using Grok-style high-toxicity synthetic conversations.
- Child-safety and anti-bullying tools, relevant to the paper's funder project "Cilter: Protecting Children Online."
- Augmentation of scarce human-labelled datasets in safety-critical NLP settings where harm detection matters more than false-positive minimization.
Industry relevance. Platform trust-and-safety teams, moderation vendors, and NLP teams with small labelled CB corpora can use SynBullying to expand training data, audit blind spots (the 37.9 Harm-F1 for authentic-trained models on GPT-4o content is a concrete example), and test robustness against LLM-generated abuse, which the paper notes can be produced by humans or bots on social media.
Future Directions
- Multilingual extension. Expanding SynBullying to additional languages to support CB detection in low-resource and cross-lingual settings.
- Cross-platform generalization. Modelling social context, including platform-specific interaction patterns and user role dynamics, to better capture real-world diversity.
- Prompt engineering refinement. Iteratively improving prompts, guided by the evaluation findings, to generate higher-quality, contextually coherent, and socially plausible conversations.
- Social-scientist validation. Having selected synthetic samples reviewed by social scientists to confirm that generated interactions represent authentic CB behaviours and realistic social dynamics.
The paper also leaves unresolved how to handle the CB types that remain hard for both LLMs and humans—Defamation and Encouragement_to_Harassment in particular—and notes that the taxonomy's "Defense" category and the additional tone_types and final_status labels collected during annotation are retained for future analysis rather than used here.
Target Audience
Researchers and practitioners working on cyberbullying detection, online harm, and abusive-language NLP; dataset creators interested in synthetic data generation and LLM-as-annotator validation; trust-and-safety and content-moderation engineers who need training or stress-test data; and computational social scientists studying adolescent online interaction. It is accessible to graduate students with basic NLP background but assumes familiarity with classification metrics and agreement statistics.
Note on the paper's own figures: the abstract lists "five dimensions" for evaluation while the introduction states "six dimensions"; the body reports five comparative dimensions (lexical/conversational, sentiment/toxicity, role distribution, harm intensity, CB type distribution).
Authors’ abstract
We introduce SynBullying, a synthetic multi-LLM conversational dataset for studying and detecting cyberbullying (CB). SynBullying provides a scalable and ethically safe alternative to human data collection by leveraging large language models (LLMs) to simulate realistic bullying interactions. The dataset offers (i) conversational structure, capturing multi-turn exchanges rather than isolated posts; (ii) context-aware annotations, where harmfulness is assessed within the conversational flow considering context, intent, and discourse dynamics; and (iii) fine-grained labeling, covering various CB categories for detailed linguistic and behavioral analysis. We evaluate SynBullying across five dimensions, including conversational structure, lexical patterns, sentiment/toxicity, role dynamics, harm intensity, and CB-type distribution. We further examine its utility by testing its performance as standalone training data and as an augmentation source for CB classification.