Skip to content
AI.info

Research

Are LLMs Empathetic to All? Investigating the Influence of Multi-Demographic Personas on a Model's Empathy

Overview Research area: Natural Language Processing / LLM fairness and empathy evaluation, at the intersection of affective computing, psychology of empathy, and demographic bias analysis. Technical l

Are LLMs Empathetic to All? Investigating the Influence of Multi-Demographic Personas on a Model's Empathy
arXiv
2510.10328
Published
2025-10-11
Authors
Ananya Malik, Nazanin Sabri, Melissa Karnaze, Mai Elsherief

AI summary

Overview

Research area: Natural Language Processing / LLM fairness and empathy evaluation, at the intersection of affective computing, psychology of empathy, and demographic bias analysis.

Technical level: Intermediate. The paper is readable for someone familiar with basic NLP evaluation ideas, but it uses causal-inference framing (Average Treatment Effect), Earth Mover's Distance, and the EPITOME empathy framework.

Scope: A quantitative and qualitative study of how four LLMs' affective (emotion-understanding) and cognitive (response-generation) empathy shifts across 315 personas built from intersecting age, gender, and culture attributes drawn from the ISEAR dataset.

What This Paper Is About

People's emotional experiences are shaped by their age, gender, and culture, so an LLM that responds empathetically to one user may respond differently, or worse, to another. The authors ask whether models show equitable empathy across user groups, whether those differences match real-world emotional patterns, and which demographic attributes the model treats as its own "neutral" default. They test this by injecting explicit demographic personas into a simulated two-turn user conversation and measuring how much the model's emotion predictions and empathic responses shift.

Key Contributions

  1. A framework for measuring empathy in LLMs across intersecting demographic attributes, rather than one attribute at a time as prior work largely did. The authors state that real-world personas are shaped by intersecting attributes and that prior work has largely overlooked this.
  2. A causal measurement design using Average Treatment Effect to compute "affective shift" (via Earth Mover's Distance over eight NRC intensity-vector emotions) and "cognitive shift" (via the EPITOME framework's Emotional Reaction, Interpretation, and Exploration scores), run in both isolation and intersection settings.
  3. A large-scale evaluation spanning 315 unique personas and 300 ISEAR emotion samples, yielding 94,500 unique model interactions across 4 LLMs: LLaMA-3-70B, GPT-4o Mini, DeepSeek-v3, and Gemini-2.0 Flash.
  4. A mixed-methods analysis combining quantitative shifts with qualitative evidence: persona recall quality, log-odds of word usage by attribute, a topic-to-attribute variance (TAV) ratio, and identification of the attributes least aligned with the model's neutral state.

Main Findings

  • Models are not uniform in empathy: The authors report substantial variation across demographic dimensions, with variation often reflecting stereotypes documented in the literature and further influenced by the type of attribute and the presence or absence of additional contextual personas.

  • Intersection attenuates and can reverse isolated effects: Adding multiple attributes at once can attenuate and reverse expected empathy patterns. For example, the male attribute is associated with higher anger intensity of approximately 0.020 above the base state, and Confucian culture with emotion intensity around 0.40 below base, but the composition shrinks the overall spread of shifts. In LLaMA-3-70B, the male anger shift moves from -0.005 in isolation to 0.003 in intersection, while female anger moves from 0.007 to -0.006.

  • Cognitive empathy ranges shrink under intersection: For LLaMA-3-70B, the cognitive age range goes from -0.206 to 0.176 in isolation to -0.0616 to 0.160 in intersection; cognitive gender from -0.613 to 0.133 to -0.512 to 0.181; cognitive culture from -0.066 to 0.073 to 0.005 to 0.108. Affective shifts shrink for age but are reported as roughly equivalent for gender and culture.

  • Specific attributes are penalized: Both affective and cognitive empathy for Confucian culture are expressed at lower levels than any other evaluated attribute across all models. The gender-queer attribute expresses higher intensities of anger across models. Overall, models tend to reduce emotion intensity relative to the base case, and GPT-4o Mini consistently lowers the cognitive empathy of its responses across nearly all personas.

  • Cognitive response quality drops in composition: For LLaMA-3-70B, exploration (EX) for the 55+ attribute moves from -0.667 in isolation to -0.003 in intersection, and the culture Emotional Reaction average drops from 0.2521 to 0.0317.

  • Weak emotion understanding overall: When outputs are unconstrained, the models show poor accuracy in the range of 0.12 to 0.18 against ground-truth labels, with mean squared error of intensity emotion vectors compared to gold labels ranging from 0.14 to 0.21. The authors note this is substantial because most intensity scores within the NRC Lexicon fall between 0.0 and 0.2.

  • Persona recall varies by model: Using cosine similarity and ROUGE-L F1 against the injected persona, DeepSeek-V3 scores 0.932 similarity (std 0.129) and 0.878 ROUGE-L (std 0.222), and Gemini 2.0 Flash scores 0.843 (std 0.144) and 0.683 (std 0.277). LLaMA-3-70B scores 0.677 (std 0.136) and 0.359 (std 0.178), and GPT-4o Mini scores 0.652 (std 0.21) and 0.514 (std 0.259).

  • Only loose alignment with real-world emotional patterns: The model assigns higher emotional intensities to the 0–17 attribute and consistently lower intensities for older age groups, matching findings that younger people are more emotionally expressive and older adults less so. Females show higher values overall across most emotions and are skewed toward positive emotions, with male personas more frequently expressing negative emotions such as anger; the authors observe this pattern in 4 out of 8 experimental settings on gender. However, the models do not consistently replicate real-world cultural variation: African-Islamic culture is reported in human baseline data to have the highest levels of anger, which is not reflected in model outputs, and the models instead assign significantly lower anger intensities to Confucian cultures.

  • No fixed neutral persona, but a biased default: Qualitative assessment of the base-state personas shows they are generic, focusing on topics and behaviors from the post and devoid of any gender, age, or culture. The authors state the model's value system is most aligned with attributes such as Protestant Europe for the anger emotion. The attributes with the maximum significant shift are the ones least aligned to the model, with 0–17 age attributes and gender-queer and Confucian culture frequently among the least aligned.

  • Responses drift toward stereotypes about the attribute: Using the topic-to-attribute variance (TAV) ratio, where a score greater than 1 implies the model is skewed toward generating responses that reflect more upon the attribute's characteristics, differences appear mainly for cultural groups, specifically Confucian, African-Islamic, and Latin-American. The DeepSeek v3 model assigns a higher ratio to the gender-queer attribute.

  • Lexical markers differ by persona: Log-odds of word usage (calculated using a Dirichlet prior) show gendered terms such as "mate" and "dude" for male, "daughter," "señora," and "gosh" for female, "attuned," "gender," and "margin" for non-binary, and "gender," "expressing," and "lgbtq" for gender-queer, along with age- and culture-specific terms including "filial" and "piety" for Confucian and "grog" and "English" for English Speaking.

Methodology in Plain English

The researchers take the ISEAR dataset, a collection of 8,000 self-reported emotional experiences from roughly 3,000 individuals, labeled with 7 emotions (anger, disgust, fear, guilt, joy, sadness, and shame). They filter out samples shorter than 10 tokens, embed the remaining sentences with SentenceTransformer's MiniLM-L6-v2 along with their gold emotion labels, and use CoreSet selection with a K-Center Greedy algorithm to pick 300 diverse samples. Samples that reveal their own emotion in the text (for example, "I feel angry...") are masked so the model cannot simply copy the answer; only 28 out of 300 samples contain a [MASK].

Each persona is assembled from 3 demographic categories: 6 age categories (0–17, 18–24, 25–34, 35–44, 45–54, 55+), 4 gender categories (male, female, non-binary, gender-queer), and 8 cultures from the Inglehart–Welzel Cultural Map, which divides 197 countries into 8 categories (Protestant Europe, English Speaking, Catholic Europe, Confucian, West and South Asia, Latin America, African-Islamic, Orthodox Europe). Each category has a "base" state in which no explicit attribute is added, so the effect of any one attribute can be isolated. Combinations across these attributes yield 315 unique persona configurations.

Conversations are simulated over 2 turns: the user states their persona ("I am a [persona]. Who am I?") and then supplies the emotional experience. For affective empathy, the model predicts the emotion (and recalls the persona, which is checked with cosine similarity and ROUGE-L). For cognitive empathy, the model generates a response, scored with the EPITOME framework's Emotional Reaction, Interpretation, and Exploration dimensions on a 0–2 scale.

Effects are quantified as shifts: the affective shift is the Earth Mover's Distance between the emotion prediction with the attribute present and absent, and the cognitive shift is the analogous difference in EPITOME scores. These are aggregated into Average Treatment Effects in two settings: isolation, where a single attribute is compared against no attribute, and intersection, where the marginal contribution of a focal attribute is measured while averaging over the other attributes present.

Why This Matters

Impact on research. The paper argues that prior work testing LLM empathy has largely used singular personas, and that this overlooks how real people are defined by intersecting attributes. It provides a reusable causal design (isolation versus intersection, using Average Treatment Effect) and a shift-based metric for both affective and cognitive empathy, and it claims to extend existing research with a more comprehensive and nuanced evaluation.

Real-world applications.

  • Mental health and healthcare chatbots, where the paper notes LLMs are increasingly deployed and where therapist-patient emotional alignment is reported to affect outcomes.
  • Emotion-aware conversational agents that must adapt to culturally grounded expressions of feeling.
  • Bias auditing of deployed models before release, using persona sweeps to find groups receiving reduced empathic quality.
  • Culturally sensitive support tools, where the authors warn that overemphasizing cultural context at the expense of emotional depth produces stereotypical understanding.

Industry relevance. Deployers choosing among the four tested model families can see that empathy behavior differs sharply by model, not just by user. The authors advocate for an alignment framework that can quantify and ensure that a model is empathetic while being emotionally intelligent and fair, arguing that current LLMs do not exhibit uniform empathetic behavior across demographic attributes.

Future Directions

  • Incorporate additional persona dimensions beyond age, gender, and culture, such as behavioral, preferential, and other lived experiences, to build more representative and complex personas.
  • Move beyond the ISEAR dataset, which is noted as not fully capturing global cultural representation or contemporary modes of emotional expression and was collected through structured surveys rather than natural conversations.
  • Study implicit persona cues, longer conversational interactions, and multimodal settings, since the current method relies on explicitly providing the persona.
  • Address prompt sensitivity and decoding variability, and investigate how to ensure personalized empathetic responses do not reinforce harmful stereotypes.

Target Audience

Researchers and practitioners in NLP fairness, affective computing, and human-AI interaction; developers and product teams building mental health, healthcare, or emotionally responsive conversational systems; and social scientists interested in how cultural and demographic emotion norms are reproduced or distorted in model outputs. Readers with a background in evaluation methodology will get the most from the causal framing, though the high-level findings are accessible without it.

Authors’ abstract

Large Language Models' (LLMs) ability to converse naturally is empowered by their ability to empathetically understand and respond to their users. However, emotional experiences are shaped by demographic and cultural contexts. This raises an important question: Can LLMs demonstrate equitable empathy across diverse user groups? We propose a framework to investigate how LLMs' cognitive and affective empathy vary across user personas defined by intersecting demographic attributes. Our study introduces a novel intersectional analysis spanning 315 unique personas, constructed from combinations of age, culture, and gender, across four LLMs. Results show that attributes profoundly shape a model's empathetic responses. Interestingly, we see that adding multiple attributes at once can attenuate and reverse expected empathy patterns. We show that they broadly reflect real-world empathetic trends, with notable misalignments for certain groups, such as those from Confucian culture. We complement our quantitative findings with qualitative insights to uncover model behaviour patterns across different demographic groups. Our findings highlight the importance of designing empathy-aware LLMs that account for demographic diversity to promote more inclusive and equitable model behaviour.

Read the original paper