Skip to content
AI.info

Research

MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes

Overview Research area: Natural Language Processing, specifically multimodal (text + image) meme analysis applied to mental health — fine-grained depressive symptom identification. Technical level: In

MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes
arXiv
2512.25015
Published
2025-12-31
Authors
Siddhant Agarwal, Adya Dhuler, Polly Ruhnke, Melvin Speisman, Md Shad Akhtar, Shweta Yadav

AI summary

Overview

  • Research area: Natural Language Processing, specifically multimodal (text + image) meme analysis applied to mental health — fine-grained depressive symptom identification.
  • Technical level: Intermediate. The clinical framing is accessible, but the method involves multi-agent LLM prompting, confidence-weighted voting, and multi-label classification metrics.
  • Scope: The paper introduces the RESTOR Ex dataset (with human and LLM-generated explanations) and MAMA-Memeia, a multi-agent, multi-aspect discussion framework grounded in Cognitive Analytic Therapy (CAT) competencies for classifying seven depressive symptoms in memes.

What This Paper Is About

Memes have become a common way for people to express depressive feelings, but most meme-analysis research has focused on harmfulness, hatefulness, and cyberbullying rather than mental health. This paper builds a resource and a method for detecting seven fine-grained depressive symptoms in memes, based on the clinically established 9-scale Patient Health Questionnaire (PHQ-9). The goal is to improve on prior work by combining multimodal LLM-generated explanations, human-written explanations, and a collaborative multi-agent debate setup guided by clinical psychology principles.

Key Contributions

  1. Novel dataset: RESTOR Ex, derived from the RESTORE dataset, adding LLM-generated explanations and human-annotated gold-label explanations, with corrected annotations and filtering of non-meme images.
  2. Novel methodology: MAMA-Memeia, a multi-agent discussion framework that adapts CAT Competencies into a multi-aspect prompting setup, described as a new state-of-the-art for the task.
  3. In-depth human evaluation: A domain-expert human evaluation of LLM-generated explanations across five aspects (fluency, relevance, figurative meaning, persuasiveness, and appeal), conducted on six model explanations plus human-annotated explanations.
  4. Comprehensive benchmarking: The framework is established as a new benchmark compared to over 30 methods, including unimodal, multimodal, single-agent LLM, and multi-aspect prompting setups.

Main Findings

  • State-of-the-art improvement: MAMA-Memeia reaches 72.73 macro-F1 and 72.45 weighted-F1, improving on the previous state of the art (Yadav et al. 2023) by 7.55% in macro-F1 and 7.78% in weighted-F1.
  • Dataset size: RESTOR Ex contains 7,096 training samples, 520 test samples, and 310 validation samples, covering labels LOI, FD, ED, SD, LSE, CP, and SH. The curation removed non-meme images, leading to a reduction of about 20% in dataset size.
  • Removed label: The 'Lack of Energy' label was removed because all of its training samples (471 samples) were automatically curated with significant issues.
  • Annotation reliability: Inter-annotator agreement for symptom labeling was 0.833 Krippendorff's Alpha using MASI distance, representing strong agreement.
  • Text matters more than image: Performance jumped sharply from unimodal image setups (ViT 34.96, ResNet 27.14, EfficientNet 25.18 macro-F1) to unimodal text setups (OCR BERT 62.02, MentalBERT 63.77, MentalBART 61.76 macro-F1).
  • Explanations beat OCR: Substituting LLM-generated explanations for OCR text as the textual modality improved BART in macro-F1 from 44.71% to 55.81%.
  • LLM explanations rival human ones: Compared to human explanations, MAMA-Memeia improved by more than 8% and Claude 3.5 Sonnet by more than 4% in weighted-F1, supporting LLM explanations as an automated, low-resource alternative.
  • Closed-source models lead: Closed-source models such as Claude 3.5 Sonnet performed significantly better than open-source models such as LLaVA 1.5 in the single-agent explanation setups.
  • Ablation: Removing aspect-specific prompting drops MAMA-Memeia's macro-F1 from 72.73 to 71.88 and weighted-F1 from 72.45 to 71.51, showing the contribution of the CAT-based multi-aspect prompts.
  • Qualitative trends: Gemini-2.0-flash tends to over-predict and produces the lengthiest explanations, while GPT-4o tends to under-predict; averaging confidence across agents helps correct both behaviors. In one example, only Claude 3.5 Sonnet correctly predicted Self-Harm and corrected the other two models across debate rounds.

Methodology in Plain English

The researchers first cleaned up an existing meme dataset (RESTORE) by re-annotating test and validation samples, removing images that were not really memes (such as quote graphics), and dropping one symptom label whose training data was unreliable. They then added two kinds of explanations for memes: human-written explanations from domain experts, and explanations generated by six multimodal LLMs (three open-source: LLaVA 1.5, LLaVA-NeXT, MiniCPM-V; three closed-source: GPT-4o, Claude 3.5 Sonnet, Gemini-2.0-flash).

For the detection method, they drew on Cognitive Analytic Therapy, a form of talking therapy that focuses on patterns of thinking, feeling, and behavior. They adapted eight CAT criteria into three knowledge aspects: Depression Knowledge (the definitions of the seven symptoms), Emotional Knowledge (the emotional state behind the meme), and Cultural Knowledge (pop culture references and figurative language). Each aspect becomes a separate prompt.

They then deploy three LLM agents (GPT-4o, Claude 3.5 Sonnet, Gemini-2.0-flash), each assigned an aspect-specific prompt, and run them through three phases: Independent Ideation, where each agent makes its own predictions, confidence estimates, and explanation; Collaborative Discussion, where agents see each other's explanations, predictions, and confidence scores over multiple rounds and can revise; and Consensus Resolution, where a confidence-weighted vote decides the final labels using a threshold to filter out low-confidence labels.

Why This Matters

  • Impact on research: It opens up the multimodal mental health domain by pairing a corrected, explanation-augmented dataset with an LLM-agent method, and shows that LLM-generated explanations can substitute for expensive human annotation at scale.
  • Real-world applications:
    • Content moderation and platform safety triage that flags memes expressing severe symptoms such as Self-Harm.
    • Mental health support communities that want to route users toward peer support or resources.
    • Automated or semi-automated annotation pipelines for building larger mental health datasets.
    • Research tooling for studying how figurative language, culture, and humor intersect with expressions of depression.
  • Industry relevance: Social media platforms, mental health technology companies, and content-moderation teams have a direct interest in scalable symptom detection, though the paper's ethics statement cautions against deploying such systems in high-stakes mental health or content-moderation settings without substantial expert human oversight, particularly for self-harm.

Future Directions

  • Open-source alternatives: The authors explicitly state that while MAMA-Memeia uses closed-source models, they look forward to developing methods for effective use of open-source LLMs, given the need for open and transparent research in sensitive domains like mental health.
  • Understanding model behavior: The qualitative analysis of over-prediction (Gemini-2.0-flash) and under-prediction (GPT-4o) is only partly explained by explanation length, and the authors note that the black-box nature of these models requires further analysis.
  • Explanation quality: The paper frames RESTOR Ex as a resource for future work comparing model-generated explanations against the human-annotated dataset.
  • Transparency and bias: The ethics statement notes that closed-source models provide limited transparency about architectures and training data, constraining interpretability, and that LLMs may introduce additional biases — both open problems.

Target Audience

Researchers and practitioners in NLP, multimodal machine learning, computational social science, and digital mental health who are interested in meme understanding, multi-agent LLM frameworks, or low-resource annotation via generated explanations. It is also relevant to clinical and psychology-informed audiences who want to see how frameworks like Cognitive Analytic Therapy can be operationalized as model prompts, and to platform safety teams evaluating the feasibility of automated symptom detection.

Authors’ abstract

Over the past years, memes have evolved from being exclusively a medium of humorous exchanges to one that allows users to express a range of emotions freely and easily. With the ever-growing utilization of memes in expressing depressive sentiments, we conduct a study on identifying depressive symptoms exhibited by memes shared by users of online social media platforms. We introduce RESTOREx as a vital resource for detecting depressive symptoms in memes on social media through the Large Language Model (LLM) generated and human-annotated explanations. We introduce MAMAMemeia, a collaborative multi-agent multi-aspect discussion framework grounded in the clinical psychology method of Cognitive Analytic Therapy (CAT) Competencies. MAMAMemeia improves upon the current state-of-the-art by 7.55% in macro-F1 and is established as the new benchmark compared to over 30 methods.

Read the original paper