Skip to content
AI.info

Research

Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation

Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation Overview Research area: Natural Language Processing, applied to mental health — specifically

Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation
arXiv
2511.01482
Published
2025-11-03
Authors
Neha Sharma, Navneet Agarwal, Kairit Sirts

AI summary

Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation

Overview

Research area: Natural Language Processing, applied to mental health — specifically automated detection of cognitive distortions (CDs) in text, plus a new evaluation methodology for comparing models trained on datasets of differing characteristics.

Technical level: Intermediate. The paper is readable without deep clinical or statistical background, but it assumes familiarity with classification metrics (F1), annotation agreement measures (Fleiss' kappa, Cohen's kappa), and transformer fine-tuning.

Scope: The paper proposes using repeated, independent LLM runs as a reliable annotation strategy for subjective text labeling, and a kappa-based effect size measure for dataset-agnostic model comparison, tested on the Therapist Q&A dataset with the MentalRoBERTa classifier.

What This Paper Is About

Automated cognitive distortion detection is difficult because the task is subjective: even expert human annotators disagree, producing noisy labels that models then learn to reproduce. This paper asks whether large language models can act as more consistent annotators than humans, by running the same text through multiple independent LLM calls and keeping only labels that recur, and it proposes a way to compare model performance fairly across datasets that differ in size and label distribution.

Key Contributions

  1. An LLM-based annotation framework. Each text is passed through five independent API calls per prompt type across four model–temperature configurations, and only labels that recur in at least 4 out of 5 runs are retained as confident annotations, treating recurrence as a signal of model certainty rather than relying on a single stochastic prediction.

  2. Statistical validation of annotation reliability. Using Fleiss' kappa per label (each of the 12 label classes treated as a separate binary task), the authors show moderate-to-substantial inter-run agreement across configurations, with GPT-4 at temperature 0.5 reaching 0.78.

  3. A dataset-agnostic evaluation methodology. Adapting the psychological concept of effect size, the paper applies Cohen's kappa to normalize a model's weighted F1 against a mathematically derived random baseline, producing a standardized score (κ_F1) that supports cross-dataset and cross-study comparison where raw F1 falls short.

  4. Expert human verification of the labels. Three psychology experts compared 101 disputed samples (where LLM and golden labels disagreed) in Label Studio; the result was inconclusive, with expert opinion evenly split between the two label sources.

Main Findings

  • High inter-run agreement for GPT-4: Fleiss' kappa averaged across the 11 CD labels was 0.78 for GPT4-0.5 on both the Ranked-Label Prompt (RLP) and Multi-Label Prompt (MLP); GPT4-0.7 scored 0.73 (RLP) and 0.71 (MLP); GPT4o-0.5 scored 0.63 (RLP) and 0.62 (MLP); GPT4o-0.7 scored 0.52 (RLP) and 0.54 (MLP).

  • Label recurrence is common: Across all configurations, at least 84% of data points had at least one label repeated four or more times. GPT-4 at temperature 0.5 showed the highest consistency, with over 90% of data points having at least one label repeated in all runs. GPT-4 was more consistent than GPT-4o, which assigned more varied label sets across runs.

  • Models trained on LLM labels beat models trained on human labels: On weighted F1 (MentalRoBERTa, averaged over five random initializations), for example with GPT4-0.7, RLP labels yielded 0.854 (binary), 0.604 (multiclass), and 0.548 (multilabel) versus 0.770, 0.391, and 0.332 for golden labels. The MLP setting had no multiclass task (marked N/A).

  • The gap holds after chance normalization: Under κ_F1, GPT4-0.7 with RLP labels on the multilabel task showed a 33.1% improvement over the random baseline, while golden labels managed only 16.1%. LLM-generated labels (RLP and MLP) consistently outperformed golden labels across all datasets and classification tasks.

  • Human verification was inconclusive: The three experts split nearly evenly between LLM-generated and golden labels. Expert 3 chose "None" in 45% of cases, versus 5% and 9% for Experts 1 and 2. Inter-expert Fleiss' kappa was 0.20 overall; pairwise agreement was κ = 0.44 between Experts 1 and 2, κ = 0.16 between Experts 1 and 3, and κ = 0.11 between Experts 2 and 3.

  • A likely cause of disagreement: In disputed examples, the user inputs often describe events, emotions, or experiences rather than explicitly stated thoughts. The authors suggest the data source itself limits what can be reliably annotated.

  • Label drift beyond the schema: Despite explicit instructions, the models sometimes generated new or modified labels outside the predefined list; these were grouped into a single "Others" category, giving 12 label classes in total (10 predefined CDs, No Distortion, and Others).

Methodology in Plain English

The researchers started from a public Therapist Q&A dataset of 2,530 user input/response pairs annotated with ten cognitive distortions by Shreevastava and Foltz (2021), whose two-annotator agreement measured 33.7% by joint probability of agreement. These original labels are called "golden labels."

Instead of asking an LLM for one answer per text, they queried GPT-4 and GPT-4o through Microsoft Azure at temperatures 0.5 and 0.7, using two prompts: a Multi-Label Prompt that allows any number of distortion labels, and a Ranked-Label Prompt that asks for one dominant distortion with an optional secondary one. Each user input was passed through five independent calls per prompt type across all four model–temperature setups, producing 40 annotations per input in total.

The reasoning is that a label appearing repeatedly across independent runs is probably picking up on a real pattern rather than sampling noise. To formalize this, they measured how often labels repeated, then computed Fleiss' kappa across runs. For the final annotation set, they kept only labels appearing in at least 4 of 5 runs — a stricter threshold than majority voting (3 of 5) — and marked inputs failing this threshold as ambiguous, removing them.

For the downstream test, the data was split 70:15:15 with fixed assignments so the same input always fell in the same split, and MentalRoBERTa was fine-tuned for binary, multiclass, and multilabel classification, trained five times with different random initializations and averaged. Because the four resulting datasets had different sizes after ambiguous examples were removed (2,369, 2,319, 2,123, and 1,888 examples versus 2,530 in the original), direct F1 comparison was not valid. The authors therefore derived a random baseline weighted F1 analytically — for a three-class case it equals a² + b² + c² — and normalized model performance with Cohen's kappa, giving κ_F1 = (F1_calculated − F1_random) / (1 − F1_random), where 0 means chance-level and 1 means perfect.

Why This Matters

Impact on research: The paper reframes annotation for subjective tasks from single-shot prediction to repeated sampling with consistency filtering, and argues that reliability must come before validity when ground truth is unattainable. It also offers a concrete way to compare models trained and tested on datasets of different sizes and label distributions, addressing a comparability problem that prior work on this dataset did not resolve. The authors note that earlier CD detection work reported weighted F1 scores in the range of 0.2–0.4, reflecting confusion between distortion categories.

Real-world applications:

  • Generating training data for mental-health text classifiers where expert annotation is expensive, slow, or hard to obtain at scale.
  • Supporting automated cognitive reframing or reappraisal tools that help users identify distorted thinking patterns in negative thoughts.
  • Screening or triage systems that flag potentially distorted thinking in online forum or social media text for follow-up.
  • Any subjective NLP labeling project — toxicity, sentiment, empathy, intent — where human annotator agreement is low and labels are expensive.

Industry relevance: The consistency-based annotation pipeline and the κ_F1 metric are both directly deployable. Organizations that build labeling pipelines can use repeated LLM sampling with a recurrence threshold to reduce annotation noise at lower cost than expert labeling, and the normalized effect-size metric gives teams a defensible way to report model improvements when their internal datasets differ in size or class balance from public benchmarks.

Future Directions

  1. Extend beyond the GPT family. The annotation experiments used only GPT-4 and GPT-4o via Azure. The authors explicitly flag that it is unclear whether Claude, Gemini, or open-source models such as LLaMA or Mistral would achieve similar annotation quality and agreement, and call for a systematic cross-family comparison.

  2. Address dataset-level limitations. The human verification results suggest many user inputs describe events and emotions rather than articulated thoughts. The authors suggest developing methodologies that ensure distorted thought patterns are adequately expressed in the text, and in a clinical setting a therapist would first probe for the underlying thought before naming a distortion.

  3. Test whether reliability implies validity. The paper positions reliability as a necessary but not sufficient step; the inconclusive expert study leaves open how validity of LLM-generated mental-health labels should be established.

  4. Apply κ_F1 outside cognitive distortion detection and outside NLP. The authors state the kappa-based effect size measure generalizes to any field where benchmarks are unstable or direct comparison is not feasible.

Target Audience

Researchers and practitioners in NLP and computational mental health who work on subjective annotation tasks, low-agreement datasets, or LLM-as-annotator pipelines. It is also relevant to clinical psychology researchers interested in whether machine-generated labels can substitute for expert annotation, and to machine learning engineers who need a principled way to compare models trained on datasets that differ in size and label distribution. Readers seeking finalized clinical validation should note that the expert verification here was inconclusive.

Authors’ abstract

Text-based automated Cognitive Distortion detection is a challenging task due to its subjective nature, with low agreement scores observed even among expert human annotators, leading to unreliable annotations. We explore the use of Large Language Models (LLMs) as consistent and reliable annotators, and propose that multiple independent LLM runs can reveal stable labeling patterns despite the inherent subjectivity of the task. Furthermore, to fairly compare models trained on datasets with different characteristics, we introduce a dataset-agnostic evaluation framework using Cohen's kappa as an effect size measure. This methodology allows for fair cross-dataset and cross-study comparisons where traditional metrics like F1 score fall short. Our results show that GPT-4 can produce consistent annotations (Fleiss's Kappa = 0.78), resulting in improved test set performance for models trained on these annotations compared to those trained on human-labeled data. Our findings suggest that LLMs can offer a scalable and internally consistent alternative for generating training data that supports strong downstream performance in subjective NLP tasks.

Read the original paper