Skip to content
AI.info

Research

Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability

Overview Research area: Natural Language Processing, specifically zero-shot stance detection (ZSSD) and interpretable machine learning. Technical level: Intermediate — the paper assumes familiarity wi

Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability
arXiv
2511.03635
Published
2025-11-05
Authors
Apoorva Upadhyaya, Wolfgang Nejdl, Marco Fisichella

AI summary

Overview

Research area: Natural Language Processing, specifically zero-shot stance detection (ZSSD) and interpretable machine learning.

Technical level: Intermediate — the paper assumes familiarity with stance detection benchmarks, information retrieval ranking, and LLM prompting, but the architecture is described at a conceptual level that a motivated non-specialist can follow.

Scope: The paper proposes IRIS, a three-stage zero-shot stance detection framework that combines LLM-generated implicit rationales (text subsequences) with linguistically grounded explicit rationales, treating stance detection as an information retrieval ranking task to gain inherent interpretability.

What This Paper Is About

Zero-shot stance detection asks a model to determine whether a post supports, opposes, or is neutral toward a target it has never seen during training. Existing approaches — contrastive learning, meta-learning, data augmentation, or LLM-based explanation generation — tend to generalize poorly or produce coarse, high-level explanations that miss word-level cues. This paper builds an interpretable system that surfaces both implicit cues (which specific snippets of a post carry stance evidence) and explicit cues (the emotional and cognitive dimensions of the author's language), and uses those as the basis for the final prediction.

Key Contributions

  1. The authors state this is the first study to frame stance detection as an information retrieval ranking task while providing inherent interpretability through both implicit and explicit reasoning.
  2. A novel interpretable ZSSD framework, IRIS (Interpretable Rationales for Stance Detection), consisting of three stages: relevance ranking, grouping and selection, and classification.
  3. A relevance ranking stage that automatically assigns each implicit rationale to favor, against, or neutral stances, plus a grouping and selection phase that picks the k most diverse rationales — removing the need for human-annotated ground truth for rationale stances.
  4. A classification phase that encodes implicit and explicit rationales, predicts a stance for each selected rationale, and determines the final label by majority voting (defaulting to neutral if no decision is reached).

Main Findings

  • Fine-tuning beats prompting for LLMs: On VAST (zero-shot targets), Mistral scored 67.09 in zero-shot, 69.76 few-shot, and 71.77 fine-tuned. On EZ (noun-phrase targets), Mistral scored 42.58, 44.06, and 53.4 respectively. Llama 3.1 scored 63.03 (zero-shot), 67.55 (few-shot), and 72.81 (fine-tuned) on VAST, and 55.13, 58.54, and 63.28 on EZ.
  • Llama 3.1 outperforms Mistral: The authors attribute this to better instruction alignment, and use Llama 3.1 to generate rationales rather than to classify stance directly.
  • LLMs are weak at picking a single best rationale: When manually annotated ground truth was compared against Llama's single best implicit rationale across 100 randomly selected correct and incorrect predictions, the F1 score was only 0.617. The authors note LLM-generated rationale quality was satisfactory when asked for all possible implicit rationales and linguistic rationales, which motivated the design shift.
  • IRIS is competitive even under limited supervision on VAST: IRIS trained on 10% of data reached 78.68 macro-F1, on 30% reached 82.22, and on 50% reached 85.56. The strongest listed baseline on VAST was Infuse at 81.43, followed by LKI-BART at 79.2 and LOT at 78.6. IRIS at 50% training data therefore exceeds all baselines listed in the provided table.
  • Per-class IRIS results on VAST (50% training data): Pro 81.15, Con 82.06, Neutral 93.47.
  • IRIS at 10% training data: Pro 73.48, Con 75.63, Neutral 86.92 — competitive with several fully trained baselines.
  • Explicit rationales are grounded in LIWC: The framework derives scores for empathy, allure, absolutist, action, concrete, agency_language, communion_language, and approach. The avoidance measure was omitted as the inverse of approach.
  • Generalizability is claimed across four datasets: VAST and EZ-STANCE for ZSSD; P-Stance for in-target and zero-shot political discourse; RFD for in-target long-form news articles — with training at 50%, 30%, and even 10% of available data. The detailed result tables for EZ, P-Stance, RFD, the ablation studies, human evaluations, and case studies are not included in the truncated content provided, so those numbers are not reported here.
  • Rating: Standard deviations for the reported metrics appear in the original tables but are cut off in the provided content.

Methodology in Plain English

The pipeline has three stages, with rationale generation feeding into the first of them.

First, an LLM (Llama 3.1) reads the post together with its target and produces two kinds of output. The explicit rationales are short written assessments of the post against LIWC-style linguistic dimensions — for example, whether the language is empathetic or absolutist, action-oriented or abstract — summarised in three to five lines. The implicit rationales are the actual word subsequences in the post that carry stance evidence. Because a single post can contain evidence for opposing positions, the system extracts all possible implicit rationales, not just one.

Second, the implicit rationales go through a relevance ranking stage. Each rationale and its target are treated as a query against three documents built from publicly available zero-shot stance datasets — one document standing in for favor, one for against, one for neutral. Stance labels are deliberately stripped from these documents, and only statements with cosine similarity below 0.05 to the training and test data are included, to avoid leakage. A pre-trained ranker (the authors selected FlagReranker over other options, and also used bge-reranker-large) produces three raw relevance scores, which are passed through softmax to give a probability distribution over favor, against, and neutral.

Third, a relevance determiner splits rationales into relevant and irrelevant groups using a threshold rule: a rationale is relevant for a stance if its score exceeds the maximum of the other two scores by more than a threshold. Then a KL-divergence based selection procedure picks k rationales that best match a target distribution reflecting how many rationales are available per subgroup, with a fallback mechanism when a subgroup runs out, so that favor/against/neutral coverage stays balanced.

Finally, the selected implicit rationales and the explicit rationales are encoded with sentence embeddings, passed through dense layers, concatenated, and fed to a softmax layer that predicts a stance for each rationale. The final stance is decided by majority vote, defaulting to neutral if no decision is made. Two losses are used: a categorical cross-entropy stance loss, and a custom "rationale usefulness reward punish" loss that rewards the ranking stage when relevant rationales lead to correct predictions and penalises it otherwise. The total loss combines them with a weight q set to 0.5.

Hyperparameters include an embedding dimension of 4096, a dense layer of 128 with ReLU, a relevance threshold of 0.3, k of 3, a reward/punish beta of 0.1, Adam at learning rate 0.0001, batch size 32, and 3 output neurons for VAST/EZ versus 2 for P-Stance/RFD. Tuning used TPE in Hyperopt for parameters and Grid Search for the loss weight.

Why This Matters

Impact on research. The paper reframes stance detection from a classification problem into a ranking problem over rationales. That reframing is what allows the model to assign stances to individual rationales without any human annotation of those rationales, which is a practical bottleneck in interpretable NLP. It also argues against the prevailing pattern of using LLMs mainly to generate post-hoc explanations or target-specific knowledge, and instead uses the LLM as a rationale extractor while the interpretability comes from the architecture itself. The authors report their model is competitive with far less training data, which matters for low-resource targets where annotated data simply does not exist.

Real-world applications:

  • Social media and platform moderation, where flagging a post as supporting or opposing a topic requires an auditable reason rather than a black-box score.
  • Political and public-opinion monitoring, where analysts need to see which phrases in a post drove an "against" classification.
  • Misinformation and claim verification, where determining whether an article supports or opposes a claim is a core sub-task.
  • Market and brand sentiment analysis on topics that did not exist when the training data was collected.

Industry relevance. The 10%-training-data result is directly relevant to companies that cannot afford large annotation campaigns for every new topic. The use of quantised 4-bit Llama and Mistral models, and off-the-shelf rankers such as FlagReranker and bge-reranker-large, means the pipeline is buildable from publicly available components rather than requiring custom model training. The reward/punish loss is also a reusable idea for any pipeline where a learned retrieval stage feeds a downstream classifier.

Future Directions

  • The provided content does not report the full EZ-STANCE, P-Stance, and RFD comparisons, the ablation studies (including the variant that uses LLM relevance scores directly in the ranking stage), or the human and automatic evaluations of rationale quality. Those are the natural places to look for where IRIS succeeds and fails.
  • The paper reports that a standard LLM asking for the single best rationale achieves only 0.617 F1 against human annotation, leaving open how much better rationale extraction could get and whether better extraction translates directly into better stance accuracy.
  • The relevance ranking stage depends on external documents constructed from benchmark datasets with a cosine similarity below 0.05 filter; how sensitive results are to that document construction, and how well it transfers to genuinely novel domains, is not resolved in the text provided.
  • The threshold of 0.3 and the selection size of k=3 are fixed hyperparameters; whether these need per-domain retuning, and how the framework scales to much longer documents such as the 416-word average RFD articles, remain open.

Target Audience

Researchers and graduate students working on stance detection, argument mining, or interpretable NLP; practitioners building topic-agnostic opinion or moderation systems who need auditable predictions; and anyone interested in using LLMs as components of a structured pipeline rather than as end-to-end classifiers. Readers looking for a general introduction to stance detection may find the paper assumes prior familiarity with zero-shot benchmarks and information retrieval ranking, but the design rationale is explained clearly enough to follow without deep expertise.

Authors’ abstract

Zero-Shot Stance Detection (ZSSD) identifies the attitude of the post toward unseen targets. Existing research using contrastive, meta-learning, or data augmentation suffers from generalizability issues or lack of coherence between text and target. Recent works leveraging large language models (LLMs) for ZSSD focus either on improving unseen target-specific knowledge or generating explanations for stance analysis. However, most of these works are limited by their over-reliance on explicit reasoning, provide coarse explanations that lack nuance, and do not explicitly model the reasoning process, making it difficult to interpret the model's predictions. To address these issues, in our study, we develop a novel interpretable ZSSD framework, IRIS. We provide an interpretable understanding of the attitude of the input towards the target implicitly based on sequences within the text (implicit rationales) and explicitly based on linguistic measures (explicit rationales). IRIS considers stance detection as an information retrieval ranking task, understanding the relevance of implicit rationales for different stances to guide the model towards correct predictions without requiring the ground-truth of rationales, thus providing inherent interpretability. In addition, explicit rationales based on communicative features help decode the emotional and cognitive dimensions of stance, offering an interpretable understanding of the author's attitude towards the given target. Extensive experiments on the benchmark datasets of VAST, EZ-STANCE, P-Stance, and RFD using 50%, 30%, and even 10% training data prove the generalizability of our model, benefiting from the proposed architecture and interpretable design.

Read the original paper