Research
Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?
Overview Research area: Natural Language Processing — fairness and bias in Large Language Models applied to zero-shot stance detection. Technical level: Intermediate. Scope: The paper measures whether

- arXiv
- 2510.20154
- Published
- 2025-10-23
- Authors
- Anthony Dubreuil, Antoine Gourru, Christine Largeron, Amine Trabelsi
AI summary
Overview
- Research area: Natural Language Processing — fairness and bias in Large Language Models applied to zero-shot stance detection.
- Technical level: Intermediate.
- Scope: The paper measures whether five LLMs change their stance predictions based on two automatically annotated sensitive attributes — African American English versus Standard American English, and text complexity measured by the Flesch-Kincaid score — across three existing stance detection datasets.
What This Paper Is About
Stance detection asks a model to label a text as being in favor of, against, or neutral toward a target such as a politician or a political issue. Because stance detection is politically sensitive, biased predictions can unfairly tie particular opinions to particular social groups. The authors investigate whether LLMs making zero-shot stance predictions fall back on stereotypes linked to how a text is written — its English dialect or its readability — rather than judging the stance itself.
Key Contributions
- The authors evaluate bias in zero-shot LLM stance detection using demographic linguistic cues on real-life data, which they state is the first work to do so (previous work by Li and Zhang (2024) used template-based synthetic data for gender bias).
- They build an automatic annotation protocol that adds two sensitive attributes to existing stance detection datasets: African American English versus Standard American English, using the model from Blodgett et al. (2016), and text complexity, using the Flesch-Kincaid readability test discretized into four groups.
- They release enhanced versions of the PStance, SCD and KE-MLM datasets that combine stance information with these sensitive attributes, along with code, at a public GitHub repository.
- They evaluate five LLMs (Mistral, Llama, Falcon, Flan and GPT-3.5) with weighted F1 for performance and Equal Opportunity as the main fairness metric, extending to Disparate Impact and Predictive Parity in the appendix.
Main Findings
- Stance detection performance hierarchy: Weighted F1 scores place Mistral and GPT-3.5 at the top, with GPT-3.5 scoring 0.787 on PStance, 0.685 on SCD and 0.695 on KE-MLM, and Mistral scoring 0.804, 0.637 and 0.671 respectively. Falcon is the weakest model, scoring 0.477 on PStance, 0.513 on SCD and 0.494 on KE-MLM.
- Llama's neutrality problem: Llama produces a high rate of neutral predictions that violate the prompt — 61.27 percent on PStance, 37.03 percent on SCD and 63.84 percent on KE-MLM — which the authors say raises concerns about its reliability for the task. Mistral also produced 65.96 percent neutral predictions on KE-MLM.
- Little bias detected on dialect: Results on the PStance dataset show very little bias based on African American English versus Standard American English. The only low-magnitude biases observed are that FLAN associates SAE more easily with being against Biden than AAE, and associates AAE more easily with being against Trump than SAE.
- Stereotype 1 — marijuana and complexity: LLMs show clear bias on the "against marijuana" target, with Equal Opportunity values reaching -0.4 for Mistral. Models are less likely to predict low complexity texts as being against marijuana, and more likely to predict high complexity texts as being against it — associating low complexity with support for marijuana and high complexity with opposition.
- Stereotype 2 — Obama and complexity: On the SCD dataset, all models except Llama exhibit bias for the Barack Obama target. Mistral, Flan and GPT-3.5 tend to predict that high complexity texts favor Obama, while Falcon shows the opposite pattern, favoring low complexity texts.
- Stereotype 3 — Biden and complexity: On the KE-MLM dataset, all models show a bias toward predicting that very high complexity texts are less likely to oppose Biden, most pronounced in Flan and GPT-3.5. Falcon again exhibits an opposite bias for very high complexity and "favor".
- Stereotype 4 — gay rights: GPT-3.5 and Llama disproportionately predict that high complexity texts are against gay rights compared to low complexity texts.
- Stereotype 5 — Falcon and partisanship: Falcon shows a similar pattern for all three political figures in the KE-MLM dataset, assigning a much higher probability that low complexity texts favor the politician regardless of party, suggesting it associates simple text with partisanship and complex text with opposition.
- Abortion shows limited bias: On the abortion topic in the SCD dataset, Equal Opportunity values are around 0, indicating limited bias with respect to text complexity.
- Overall bias ranking: Averaging the absolute value of Equal Opportunity, Falcon shows the maximum bias overall (0.20 for complexity, 0.07 for A/SAE), followed by Mistral (0.12, 0.09), Flan (0.12, 0.08), Llama (0.07, 0.08) and GPT-3.5 (0.08, 0.04).
- Dataset imbalance after annotation: The PStance dataset contained 20,575 SAE tweets against only 339 AAE tweets before balancing; the study downsamples the majority group to 339 per group. SCD was balanced to 262 texts per complexity class and KE-MLM to 160 per class.
Methodology in Plain English
The authors started from three existing stance detection datasets. PStance covers three American political figures (Bernie Sanders, Joe Biden, Donald Trump) and was used for the dialect experiment. SCD covers four debate themes (Abortion, Gay rights, Marijuana, Barack Obama) and KE-MLM covers two political figures (Donald Trump, Joe Biden); both were used for the text complexity experiment. PStance was not used for complexity because it was heavily over-represented in medium and low complexity levels, and SCD and KE-MLM contained an insignificant proportion of AAE texts, making the dialect study inapplicable to them.
Because no existing stance detection dataset contains author-level sensitive attribute annotations, the authors annotated the texts automatically. For dialect, they used the Blodgett et al. (2016) model, which returns probabilities for four English varieties labeled "African-American", "Hispanic", "Asian" and "Standard", and assigned each text the highest-probability label. For complexity, they computed the Flesch-Kincaid score using average sentence length and average syllables per word, then discretized it into Easy (score of 80 or above), Medium (60 to below 80), Difficult (30 to below 60) and Very difficult (below 30).
They then balanced every dataset so each sensitive group had an equal number of texts, and each group had the same proportion of favorable and unfavorable posts, accepting smaller datasets to avoid class imbalance effects. Each model was prompted zero-shot using the Context Analyze prompt (chosen over Zero-shot Chain-of-Thought because preliminary experiments gave similar results and Context Analyze was much faster), asking the model to respond with a single word, FAVOR or AGAINST. Models tested were GPT-3.5-turbo-0125, Llama3-8B-Instruct, Mistral-7B-Instruct-v0.2, Falcon-7b-instruct and FLAN-T5-large.
Fairness results were computed by averaging metrics over 1000 balanced samples randomly drawn from each dataset, with the same samples used across all five models. Equal Opportunity compares the probability of predicting label 1 for one group versus its counterpart among examples whose true label is 1, and it was reported both with "favor" as label 1 and with "against" as label 1.
Why This Matters
Impact on research: The paper argues that bias evaluation in stance detection has been largely overlooked, partly because datasets rarely include sensitive attribute annotations, and that its annotation protocol could be generalized to other group categorizations such as gender. It provides a released dataset and metric framework that other researchers can build on, and it points toward debiasing approaches such as fairness-aware prompting or calibration and counterfactual causal modeling.
Real-world applications:
- Social media moderation pipelines that automatically flag or rank political posts, where dialect- or readability-driven errors could suppress or amplify specific groups' speech.
- Political analysis and voter-targeting tools that infer users' political orientation from their posts, where biased inference risks mischaracterizing individuals and communities.
- Deployment of LLM assistants that summarize or classify public opinion, where stereotypes about writing style could distort the reported distribution of views.
- Auditing and compliance work in organizations that must demonstrate that their NLP systems do not discriminate against protected groups.
Industry relevance: Because the models showing the largest overall bias are Falcon and Mistral while GPT-3.5 shows the smallest, the authors suggest the differences likely stem from pretraining corpus composition or instruction tuning data rather than architecture, since the models share similar architectures. For companies choosing which model to deploy for sensitive classification tasks, this is a practical selection signal, and it highlights that instruction-tuned open models vary widely in fairness even when their task performance is comparable.
Future Directions
- Larger datasets: the balancing procedure produced relatively small sample sizes, and the authors state that a study with a larger dataset would be beneficial.
- Larger models: the Mistral and Llama versions used have a limited number of parameters, and the authors note larger variants may perform better on stance detection and may reveal additional biases.
- Combining static and LLM-based metrics for automatic group categorization, which the authors call a promising line of research.
- Applying debiasing techniques such as fairness-aware prompting, calibration, or counterfactual inference to isolate the contribution of sensitive attributes, and generalizing the protocol to other group categorizations such as gender.
Target Audience
Researchers and practitioners working on fairness and bias in NLP, especially those evaluating or deploying LLMs for stance detection, political text analysis, or other identity-sensitive classification tasks. It is also relevant to dataset builders who need methods for annotating sensitive attributes on real-world text, and to industry teams selecting between instruction-tuned LLMs for socially sensitive deployments. Readers should be comfortable with classification metrics and fairness definitions such as Equal Opportunity, though the paper explains these in the text.
Authors’ abstract
Large Language Models inherit stereotypes from their pretraining data, leading to biased behavior toward certain social groups in many Natural Language Processing tasks, such as hateful speech detection or sentiment analysis. Surprisingly, the evaluation of this kind of bias in stance detection methods has been largely overlooked by the community. Stance Detection involves labeling a statement as being against, in favor, or neutral towards a specific target and is among the most sensitive NLP tasks, as it often relates to political leanings. In this paper, we focus on the bias of Large Language Models when performing stance detection in a zero-shot setting. We automatically annotate posts in pre-existing stance detection datasets with two attributes: dialect or vernacular of a specific group and text complexity/readability, to investigate whether these attributes influence the model's stance detection decisions. Our results show that LLMs exhibit significant stereotypes in stance detection tasks, such as incorrectly associating pro-marijuana views with low text complexity and African American dialect with opposition to Donald Trump.