Research
IndRegBias: A Dataset for Studying Indian Regional Biases in English and Code-Mixed Social Media Comments
Overview Research area: Natural Language Processing, specifically social bias and fairness in language models, dataset construction, and evaluation of Large Language Models (LLMs) and Indic Language M

- arXiv
- 2601.06477
- Published
- 2026-01-10
- Authors
- Debasmita Panda, Akash Anil, Neelesh Kumar Shukla
AI summary
Overview
Research area: Natural Language Processing, specifically social bias and fairness in language models, dataset construction, and evaluation of Large Language Models (LLMs) and Indic Language Models (ILMs) on Indian regional stereotypes expressed in English and code-mixed social media text.
Technical level: Intermediate. The paper assumes familiarity with classification metrics (accuracy, precision, F1), zero-shot and few-shot prompting, and parameter-efficient fine-tuning, but the dataset construction and annotation logic are explained in accessible terms.
Scope in one sentence: The paper introduces IndRegBias, a 25,000-comment dataset of Reddit and YouTube comments annotated for whether they contain Indian regional bias, how severe that bias is, and which Indian region is targeted, and benchmarks 11 open-source LLMs/ILMs on the task using zero-shot, few-shot, and fine-tuning setups.
What This Paper Is About
Regional bias — stereotypes or prejudice aimed at people because of the Indian state or region they come from — is common in Indian social media, but existing bias datasets for the Indian context cover race, gender, religion, and economics more than region, and the ones that touch on region are mostly written in English only. The authors' goal is to build a real-world, multilingual, code-mixed benchmark of Indian regional bias drawn from how users actually write online, and then test whether current language models can detect that bias and judge how severe it is.
Key Contributions
- A new dataset, IndRegBias: 25,000 user comments scraped from Reddit (via praw) and YouTube (via googleapiclient) from threads and videos discussing regional tensions in India, covering 36 regional boundaries defined by the States Reorganization Act and subsequent amendments.
- A multilevel annotation policy: A three-task labeling scheme assigning (i) a binary Regional Bias / Non-Regional Bias label, (ii) a severity score (Mild = 1, Moderate = 2, Severe = 3) for comments labeled as biased, and (iii) the target region from a predefined list covering broad regions (North-India, NorthEast-India, Central India, etc.) and 30+ specific states/union territories.
- A human-validated annotation with reported agreement: Two groups of three university students each (annotators knowing at least three Indian languages) labeled the same data; inter-annotator agreement was Cohen's Kappa of 0.91 for the binary task and 0.83 for the combined multi-class task, with disagreements resolved through discussion including one neutral annotator.
- A systematic benchmark of LLMs and ILMs: Evaluation of open-source and Indic models under zero-shot, few-shot, and fine-tuning (LoRA-based PEFT) conditions for both bias detection and severity classification, plus qualitative error analysis including model safety refusals.
Main Findings
- The dataset is roughly balanced on bias but skewed on geography: Of the 25,000 comments, 13,015 (52.1%) are regional bias and 11,985 (47.9%) are non-regional bias. Among biased comments, approximately 29% are Mild, 51% Moderate, and 20% Severe.
- Region distribution is uneven: Approximately 40% of biased comments target South India, 26% North India, 16% East India, 10% West India, 5% NorthEast India, and 2% Central India. Raw region counts are South-India 5,820; North-India 3,844; East-India 2,307; West-India 1,508; NorthEast-India 827; Central-India 339.
- State-level bias is long-tailed: Kerala, Goa, West Bengal, Karnataka, Bihar, and Gujarat are targeted approximately 60% of the time. Individual counts include Kerala 1,857, Goa 1,095, West Bengal 1,091, Karnataka 887, Bihar 871, Gujarat 832, Tamil Nadu 614, Uttar Pradesh 599, Maharashtra 438, Punjab 412, and as few as 12 and 16 for Tripura and Meghalaya.
- Zero-shot binary detection is weak to moderate: The best zero-shot results in Table 5 come from Qwen3-8B (accuracy 0.74, precision 0.71, F1 0.78) and Krutrim-2 (0.73, 0.68, 0.78). Phi-4-Mini scored an F1 of 0.00 despite 0.48 accuracy and 0.60 precision. The paper attributes Qwen3's strength to multilingual pre-training covering 119 languages, and Krutrim-2's to its targeted pre-training in 13 Indian languages.
- Few-shot gains are mixed: For Qwen3-8B, 50 regional-bias examples (Exp-1) raised F1 from 0.77 to 0.82; a balanced 25R/25N set (Exp-2) reached the highest precision (0.85) but dropped F1 to 0.75; a 30R/20N set (Exp-3) reached precision 0.80 and F1 0.81.
- Fine-tuning is the largest improvement: For Qwen3-8B, precision went from 0.692 zero-shot to 0.902 fine-tuned and F1 from 0.772 to 0.904. For Qwen3-32B, precision went from 0.790 to 0.900 and F1 from 0.724 to 0.902.
- Zero-shot severity classification is difficult: Mistral-7B-v0.3 achieved the highest precision for the Mild and Moderate classes (0.64) and the best F1 for Severe (0.45). Qwen3-8B performed best on Severe precision (0.44) and Moderate F1 (0.66); Qwen3-32B lagged in the Mild and Severe categories.
- Fine-tuning also helps severity: Fine-tuned Mistral-7B-v0.3 improved Moderate F1 from 0.61 to 0.72 and Severe precision from 0.36 to 0.59, with Severe F1 rising from 0.45 to 0.53, though Mild precision fell from 0.79 to 0.76.
- Three themes dominate severe bias: Qualitative analysis groups severe comments into Crime and Corruption (Delhi, Uttar Pradesh, Tripura, J&K, Goa), Social Intolerance (Delhi, Mizoram, Karnataka, Haryana, Uttar Pradesh, J&K, Manipur), and Socio-Economic Failure (Bihar, Uttar Pradesh, Jharkhand), with roughly similar counts across the three themes (for example Delhi at 33 comments, Jharkhand 37, Bihar 30).
- Safety alignment hurts measured performance: Highly aligned models such as Gemini-2.5-Pro, Llama-3.2, and Sarvam-M frequently refuse to classify slurs or derogatory stereotypes, returning neutral output or refusal strings, which lowers their recall.
- Proprietary comparison: GPT-4o and Llama 4 Scout were evaluated on a 5,000-comment snapshot, and the paper reports that open-source models such as Qwen achieved similar performance.
Methodology in Plain English
The authors first chose subreddit pages and YouTube channels that discuss stereotypes, biases, and identity issues tied to particular Indian states. They pulled comments with the official PRAW and Google API client tools, removed spam and irrelevant material, and ended with 25,000 comments to annotate. Because the source communities mix English with Hindi, Bengali, Malayalam, Marathi and other languages, the corpus naturally contains transliterated and code-mixed text.
Two three-person teams of university students annotated the same comments in three steps: is this comment a regional bias or not; if yes, how severe is it (Mild, Moderate, Severe, with positive stereotypes explicitly mapped to Mild); and which region or state is being targeted. Agreement was measured with Cohen's Kappa, and disagreements were talked through.
For the model evaluation, the authors first checked the models with no examples at all (zero-shot) and with a small number of worked examples (few-shot), using chain-of-thought prompts. The prompts instruct the model to identify regional references, look for generalizations, assess tone, and output a final classification. For few-shot work they used Qwen3-8B rather than Qwen3-32B because it is competitive but cheaper to run, and they restricted the number of support examples because of GPU cost, context-window limits (around 2,048 tokens), and data sparsity.
For fine-tuning they used LoRA adapters that keep the base weights frozen and train only small add-on parameters. The data was split 70% train, 10% validation, and 20% test with 5-fold stratified cross-validation. Binary bias detection used instruction-based supervised fine-tuning (10 epochs, learning rate 2e-4, LoRA rank 16, alpha 32), while severity prediction used classification-based fine-tuning that replaces the generative layer with a classification head (Mistral-7B-Instruct-v0.3, 5 epochs, learning rate 2e-5), with a WeightedRandomSampler to offset class imbalance.
Why This Matters
Impact on research: The paper argues that regional bias is under-studied relative to gender, race, and economic bias, and that existing Indian bias resources such as SeeGULL and IndiBias are largely English-based or translated, which may lose the local character of the bias. IndRegBias provides a naturally occurring, code-mixed, severity-annotated resource that shows off-the-shelf models — even strong ones — struggle with this task unless they are adapted on domain-specific data.
Real-world applications:
- Content moderation systems for Indian-language and code-mixed social media that need to flag region-targeted abuse rather than only generic hate speech.
- Safety alignment and red-teaming of LLMs deployed to Indian users, where rigid refusal behavior currently blocks legitimate analysis of biased text.
- Election-period and communal-tension monitoring by platforms and civil society groups tracking regional hostility online.
- Culturally aware evaluation suites for Indic language model development, giving builders a concrete benchmark for region-aware performance.
Industry relevance: The gap between zero-shot performance and fine-tuned performance (for example, F1 rising from 0.772 to 0.904 for Qwen3-8B) suggests that deploying a general-purpose model without adaptation is insufficient for regional bias detection, and that modest LoRA fine-tuning yields large gains at low computational cost. The finding that safety-aligned models refuse to classify slurs also matters for teams that need moderation tooling to analyze, not generate, offensive language.
Future Directions
- Fix the geographic imbalance. The dataset over-represents South India (approximately 40%) and North India (approximately 26%) relative to Northeast India (approximately 5%) and Central India (approximately 2%); the authors identify this as something that needs to be addressed, and attribute it to media marginalization and possible geographic or digital barriers.
- Reduce subjectivity in severity labels. The line between Mild and Moderate can shift with an annotator's cultural background, despite a Kappa of 0.83, so refining severity definitions or collecting more annotator perspectives is an open problem.
- Separate refusal from inability. Because aligned models often decline to label slurs, future evaluations could disentangle genuine failure to recognize bias from safety-triggered refusal.
- Extend coverage beyond English and code-mixed comments. The paper notes that prior resources were curated in English and may lose local bias characteristics, leaving room for broader multilingual and transliterated resources, and for testing whether larger or more multilingual pre-training (as with Qwen's 119 languages) closes the remaining gap.
Target Audience
Researchers and practitioners working on NLP fairness and bias, Indic language technologies, and content moderation for Indian social media will get the most from this paper. It is also useful for dataset builders interested in multilevel annotation schemes and inter-annotator agreement practices, and for LLM engineers who want evidence on how zero-shot, few-shot, and LoRA fine-tuning compare on a culturally nuanced classification task. Readers should heed the paper's own warning that the dataset contains offensive examples of regional bias.
Authors’ abstract
Warning: This paper consists of examples representing regional biases in Indian regions that might be offensive towards a particular region. While social biases corresponding to gender, race, socio-economic conditions, etc., have been extensively studied in the major applications of Natural Language Processing (NLP), biases corresponding to regions have garnered less attention. This is mainly because of (i) difficulty in the extraction of regional bias datasets, (ii) disagreements in annotation due to inherent human biases, and (iii) regional biases being studied in combination with other types of social biases and often being under-represented. This paper focuses on creating a dataset IndRegBias, consisting of regional biases in an Indian context reflected in users' comments on popular social media platforms, namely Reddit and YouTube. We carefully selected 25,000 comments appearing on various threads in Reddit and videos on YouTube discussing trending topics on regional issues in India. Furthermore, we propose a multilevel annotation strategy to annotate the comments describing the severity of regional biased statements. To detect the presence of regional bias and its severity in IndRegBias, we evaluate open-source Large Language Models (LLMs) and Indic Language Models (ILMs) using zero-shot, few-shot, and fine-tuning strategies. We observe that zero-shot and few-shot approaches show lower accuracy in detecting regional biases and severity in the majority of the LLMs and ILMs. However, the fine-tuning approach significantly enhances the performance of the LLM in detecting Indian regional bias along with its severity.