Skip to content
AI.info

Research

LLMs vs. Traditional Sentiment Tools in Psychology: An Evaluation on Belgian-Dutch Narratives

Overview Research area: Natural language processing for sentiment/emotion analysis, intersecting with computational linguistics and psychology, with a focus on a low-resource language variant (Flemish

LLMs vs. Traditional Sentiment Tools in Psychology: An Evaluation on Belgian-Dutch Narratives
arXiv
2511.07641
Published
2025-11-10
Authors
Ratna Kandala, Katie Hoemann

AI summary

Overview

  • Research area: Natural language processing for sentiment/emotion analysis, intersecting with computational linguistics and psychology, with a focus on a low-resource language variant (Flemish, the Belgian variant of Dutch).
  • Technical level: Intermediate — the paper assumes familiarity with lexicon-based sentiment tools, large language models, and model fine-tuning, though the core comparison is easy to follow.
  • Scope: The paper benchmarks three Dutch-specific LLMs against two traditional lexicon-based tools (LIWC and Pattern) for predicting emotional valence in spontaneous Belgian-Dutch personal narratives.

Note: this summary is based solely on the paper's abstract; metric details, statistical tests, and per-model comparisons beyond the overall ranking are not available in it.

What This Paper Is About

Sentiment tools have traditionally relied on curated word lists, which can miss context and nuance, while large language models are often assumed to handle context better. The authors test that assumption in a demanding setting: predicting how positively or negatively people felt, based on their own spontaneous written accounts of everyday experience, in a low-resource Dutch variant. The goal is to find out whether Dutch-tuned LLMs actually outperform established lexicon-based instruments on this kind of real-world emotional language.

Key Contributions

  1. A direct head-to-head evaluation of three Dutch-specific LLMs (ChocoLlama-8B-Instruct, Reynaerde-7B-chat, and GEITje-7B-ultra) against two traditional lexicon-based tools (LIWC and Pattern) for valence prediction.
  2. A dataset of approximately 25,000 spontaneous textual responses from 102 Dutch-speaking participants, each writing about current experiences and providing a self-assessed valence rating on a -50 to +50 scale.
  3. The empirical finding that the Dutch-tuned LLMs performed worse than traditional methods on this task, with Pattern achieving the best performance.
  4. An argument that evaluating emotion language in low-resource language variants requires culturally and linguistically tailored frameworks, and a challenge to whether current LLM fine-tuning approaches capture everyday emotional expression.

Main Findings

  • Traditional tools beat Dutch-tuned LLMs: Despite their architectural sophistication, the three Dutch-specific LLMs underperformed LIWC and Pattern on valence prediction.
  • Pattern was the strongest performer: Among the tools compared, the lexicon-based Pattern tool showed superior performance.
  • Spontaneous narratives are hard: The authors frame the result as evidence of how difficult it is to capture emotional valence in natural, self-reported, real-world language rather than curated or artificial text.
  • Fine-tuning may not be enough: The results raise the question of whether current fine-tuning approaches for LLMs adequately address nuanced emotional expression.
  • Low-resource variants need dedicated evaluation: The paper stresses the need for culturally and linguistically tailored evaluation frameworks for language variants such as Flemish.

Methodology in Plain English

The researchers collected around 25,000 short pieces of text written by 102 Dutch-speaking participants, each describing their current experiences. Alongside each narrative, the participants rated how they felt on a scale running from -50 (very negative) to +50 (very positive).

The researchers then asked several automated systems to predict those valence ratings from the text alone. Three of the systems were language models specifically adapted for Dutch, and two were established dictionary-based tools that score text using predefined word lists. By comparing how closely each system's predictions matched the participants' own ratings, the authors could see whether the newer, context-sensitive language models actually did better than the older, word-list approaches. The abstract does not report the specific accuracy or correlation figures used to rank the systems.

Why This Matters

Impact on research: The result pushes back on a widespread assumption that LLMs are automatically superior for sentiment and emotion tasks. It suggests that for low-resource language variants, and for spontaneous personal narratives specifically, well-established lexicon tools remain competitive or better — which has direct implications for how emotion researchers choose instruments and validate them.

Real-world applications:

  • Psychological and well-being research that uses automated coding of diary entries, experience sampling text, or open-ended survey responses.
  • Mental health and clinical monitoring tools that infer affect from what people write, particularly in languages and dialects with less training data.
  • Cross-cultural and multilingual product or social media analytics, where teams must decide whether to buy or build LLM-based sentiment pipelines.
  • Public health or social science surveys using free-text responses that need scalable, validated valence scoring.

Industry relevance: Teams building sentiment or emotion features for Dutch-speaking and other lower-resource markets get a clear signal that model size and language-specific tuning do not guarantee better performance, and that benchmark choice matters. It argues for evaluating tools on real user-generated narrative text rather than on curated benchmark sets before deploying them.

Future Directions

  • Develop culturally and linguistically tailored evaluation frameworks for low-resource language variants such as Flemish.
  • Investigate why Dutch-tuned LLMs lag behind lexicon tools on spontaneous narratives — for example, whether fine-tuning data, prompt design, or task framing is the limiting factor.
  • Explore improved fine-tuning or adaptation methods that better capture subtle, everyday emotional expression rather than explicit sentiment.
  • Compare and potentially combine lexicon-based and LLM approaches, since the two families appear to have complementary strengths.

Target Audience

Researchers in NLP and affective computing working on sentiment and emotion analysis; psychologists and social scientists who use automated text analysis on self-report or diary data; and practitioners building sentiment tools for Dutch or other low-resource language variants who need evidence about which method to trust.

Authors’ abstract

Understanding emotional nuances in everyday language is crucial for computational linguistics and emotion research. While traditional lexicon-based tools like LIWC and Pattern have served as foundational instruments, Large Language Models (LLMs) promise enhanced context understanding. We evaluated three Dutch-specific LLMs (ChocoLlama-8B-Instruct, Reynaerde-7B-chat, and GEITje-7B-ultra) against LIWC and Pattern for valence prediction in Flemish, a low-resource language variant. Our dataset comprised approximately 25000 spontaneous textual responses from 102 Dutch-speaking participants, each providing narratives about their current experiences with self-assessed valence ratings (-50 to +50). Surprisingly, despite architectural advancements, the Dutch-tuned LLMs underperformed compared to traditional methods, with Pattern showing superior performance. These findings challenge assumptions about LLM superiority in sentiment analysis tasks and highlight the complexity of capturing emotional valence in spontaneous, real-world narratives. Our results underscore the need for developing culturally and linguistically tailored evaluation frameworks for low-resource language variants, while questioning whether current LLM fine-tuning approaches adequately address the nuanced emotional expressions found in everyday language use.

Read the original paper