Research
Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
Overview Research area: Natural Language Processing / multimodal emotion analysis, specifically Meme Emotion Understanding (MEU) — classifying the emotional intent behind image-plus-text memes. Techni

- arXiv
- 2511.11126
- Published
- 2025-11-14
- Authors
- Yi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei, Wenzhi Chen
AI summary
Overview
- Research area: Natural Language Processing / multimodal emotion analysis, specifically Meme Emotion Understanding (MEU) — classifying the emotional intent behind image-plus-text memes.
- Technical level: Intermediate. The core ideas (prompting large models for richer text, then fusing text and image features) are accessible, but the dual-stage fusion design assumes some familiarity with cross-attention and multimodal encoders.
- Scope: The paper proposes MemoDetector, a framework that uses a Multimodal Large Language Model (MLLM) to generate four levels of textual interpretation for each meme, then feeds those interpretations plus the original image into a small classifier with a two-stage fusion mechanism.
What This Paper Is About
Memes carry emotion through the interaction of an image and its caption, and that emotion is often implicit — sarcasm, metaphor, or a reference that requires outside knowledge. Existing systems for Meme Emotion Understanding tend to fuse the two modalities in one simple step and to ignore background knowledge, so they miss these subtle cues. This paper's goal is to enrich the text side of a meme with model-generated interpretation and to fuse image and text twice, in a shallow pass and then a deeper bidirectional pass, to classify the emotion more accurately.
Key Contributions
-
Problem framing and framework. The authors identify two core challenges in MEU — a lack of fine-grained multimodal fusion strategies, and insufficient mining of memes' implicit meanings and background knowledge — and propose MemoDetector, which they state is the first to incorporate contextual background knowledge alongside a fine-grained modal fusion strategy for this task.
-
Four-step textual enhancement module. A Chain-of-Thought (CoT) prompt sequence with four steps: Image Description (ID), Text Meaning (TM), Combined Implicit Meaning (CIM), and Context Analysis (CA). These map onto three hierarchical levels inspired by human cognition — shallow perception, deep interpretation, and associative reasoning.
-
Dual-stage modal fusion. Stage 1 concatenates raw image patches with original text tokens to build an enriched visual representation. Stage 2 applies bidirectional cross-attention between that representation and the four enhanced text representations, so each modality can refine the other, before concatenating mean-pooled vectors for classification.
-
State-of-the-art results and analysis. Experiments on MET-MEME and MOOD, plus ablation studies and three targeted analyses (four-step vs. direct enhancement, MLLM scale and usage paradigm, and second-stage fusion strategy choice). Code is released at https://github.com/singing-cat/MemoDetector.
Main Findings
-
Best results on both benchmarks. MemoDetector achieves 49.57 accuracy and 45.33 Macro-F1 on MET-MEME, and 83.52 accuracy and 82.65 Macro-F1 on MOOD. The abstract reports F1 gains of 4.3% on MET-MEME and 3.4% on MOOD; the introduction reports accuracy gains of 4.17% and 4.04% respectively.
-
Gains over the strongest baselines. Against Early Fusion on MET-MEME, MemoDetector improves by 5.3% accuracy, 6.14% precision, 5.15% recall, and 5.39% Macro-F1. Against ALFRED on MOOD, it improves by 4.42% accuracy, 2.96% precision, 1.97% recall, and 3.4% Macro-F1.
-
Multi-stage fusion matters more than any single enhancement step. Removing the dual-stage fusion module (replaced with simple concatenation) drops Macro-F1 from 45.33 to 43.57 on MET-MEME, a 1.76% decrease, and from 82.65 to 81.39 on MOOD.
-
Removing Text Meaning hurts the most. Ablating the TM step produces the largest degradation among the four text-enhancement steps: a 1.55% F1 drop on MET-MEME and 2.11% on MOOD. All four steps contribute positively, including both the shallow ones (ID, TM) and the deeper ones (CIM, CA).
-
The four-step strategy beats a one-shot prompt. A comparison against "direct textual enhancement," where the MLLM infers the emotion and gives a single-step explanation, shows the four-step version outperforming it across all metrics on both datasets.
-
Larger MLLMs and the small-model paradigm help. Using QwenVL-32B for enhancement instead of QwenVL-7B gives better final results (49.57/45.33 vs. 45.60/41.22 on MET-MEME; 83.52/82.65 vs. 82.70/81.72 on MOOD). Chain-of-thought prompting helped QwenVL-7B (34.79 to 40.77 accuracy on MET-MEME) but hurt QwenVL-32B (38.58 to 38.18), which the authors attribute to larger models already having sufficient reasoning capacity.
-
Bidirectional cross-attention outperforms alternatives. Replacing the second-stage module with add, concatenate, or plain cross-attention consistently degrades performance on both datasets.
-
MLLM zero-shot performance is weak. Despite having the most parameters, Qwen2.5-VL-7B, Qwen2.5-VL-32B, and GPT-4.1 in zero-shot mode fall short on MET-MEME (for example, GPT-4.1 at 41.61 accuracy and 33.08 Macro-F1), while fine-tuned Qwen2.5-VL-7B reaches 45.40 accuracy and 41.03 Macro-F1, surpassing GPT-4.1 on that dataset.
Methodology in Plain English
The system works in two halves — one large model that explains, one small model that decides.
First, a Multimodal Large Language Model (Qwen2.5-VL-32B by default) is asked four sequential questions about each meme, each building on the last: describe what is visible while ignoring the text; analyze the meaning, tone, and rhetoric of the text; state what the image and text together are likely meant to convey; and suggest a real-world situation in which someone would use this meme. The answers become an "enhanced text" that captures metaphor, cross-modal meaning, and context that the raw caption lacks.
Second, a small classifier does the actual emotion prediction. The meme image is encoded with ViT and the text with XLM-R, which was chosen so the system can handle multiple languages. Fusion happens in two stages. Stage 1 concatenates the image patches with the original caption's tokens, treating the text tokens as extra "pseudo-patches" so the visual representation already carries some textual signal. Stage 2 concatenates the four enhanced text representations and runs bidirectional cross-attention between the enriched visual features and the enriched textual features, so each side attends to the other. The two resulting vectors are mean-pooled, concatenated, and passed through a linear layer and softmax to produce the emotion label. Training minimizes average cross-entropy loss, and all reported scores are averaged over 5 runs with random seeds.
Why This Matters
Impact on research: The paper argues that bigger is not automatically better for this task — zero-shot MLLMs underperform smaller, purpose-built multimodal classifiers. It offers a middle path in which a large model's knowledge is distilled into a small, cheap classifier through text rather than fine-tuning, and it shows that how you structure the prompting (four steps vs. one) matters as much as which model you use. It also contributes a documented decomposition of what each fusion and enhancement component buys, with ablations on two datasets.
Real-world applications:
- Content moderation and platform safety, where sarcastic or coded memes evade keyword- and sentiment-based filters.
- Brand and public-opinion monitoring, since memes are a common vehicle for sentiment toward products, figures, and events.
- Mental health and crisis detection in social media, where distress is often expressed indirectly through humor or metaphor.
- Recommendation and feed ranking systems that need to know how users actually feel about the content they post or share.
Industry relevance: The design avoids the impractical cost of fine-tuning a 32B-parameter model per task — the large model is used as an offline or cached knowledge source, while the deployed classifier remains small. That is a practical architecture for platforms processing high volumes of user content, and the released code lowers the barrier to reproducing it.
Future Directions
- Does the approach generalize beyond emotion? The framework is presented for MEU specifically; whether the same four-step enhancement and dual-stage fusion transfers to hate-speech detection, sarcasm detection, or multilingual moderation is not tested here.
- Cost, latency, and caching are not reported. Running a 32B MLLM for four prompts per meme is expensive, and the paper reports no inference-time or compute measurements. Whether enhanced texts can be generated once and reused, or generated at lower cost, is an open question.
- Why does CoT hurt larger models? The authors offer an explanation (larger models may already reason sufficiently and step-by-step constraints may inhibit them) but do not test it directly; this is a natural follow-up experiment.
- Dataset coverage. Evaluation is limited to MET-MEME (4,000 English and 6,045 Chinese memes, seven emotion classes) and MOOD (10,004 English memes, six emotions). Behavior on other languages, other emotion taxonomies, or on memes requiring very recent cultural context is unknown.
Target Audience
Researchers and graduate students working on multimodal emotion analysis, meme understanding, or affect detection in social media; practitioners building content moderation or social listening systems who need a small, deployable classifier backed by a larger model's world knowledge; and anyone studying how to transfer knowledge from Multimodal LLMs into lightweight task-specific models.
Authors’ abstract
With the rapid rise of social media and Internet culture, memes have become a popular medium for expressing emotional tendencies. This has sparked growing interest in Meme Emotion Understanding (MEU), which aims to classify the emotional intent behind memes by leveraging their multimodal contents. While existing efforts have achieved promising results, two major challenges remain: (1) a lack of fine-grained multimodal fusion strategies, and (2) insufficient mining of memes' implicit meanings and background knowledge. To address these challenges, we propose MemoDetector, a novel framework for advancing MEU. First, we introduce a four-step textual enhancement module that utilizes the rich knowledge and reasoning capabilities of Multimodal Large Language Models (MLLMs) to progressively infer and extract implicit and contextual insights from memes. These enhanced texts significantly enrich the original meme contents and provide valuable guidance for downstream classification. Next, we design a dual-stage modal fusion strategy: the first stage performs shallow fusion on raw meme image and text, while the second stage deeply integrates the enhanced visual and textual features. This hierarchical fusion enables the model to better capture nuanced cross-modal emotional cues. Experiments on two datasets, MET-MEME and MOOD, demonstrate that our method consistently outperforms state-of-the-art baselines. Specifically, MemoDetector improves F1 scores by 4.3\% on MET-MEME and 3.4\% on MOOD. Further ablation studies and in-depth analyses validate the effectiveness and robustness of our approach, highlighting its strong potential for advancing MEU. Our code is available at https://github.com/singing-cat/MemoDetector.