Research
CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension
Overview Research area: Multimodal natural language processing / vision-language modeling, specifically computational humor (satire, sarcasm, memes) and prompting-based reasoning with large vision-lan

- arXiv
- 2608.23172
- Published
- 2026-08-24
- Authors
- Abhilash Nandy, Rahul Seetharaman, Aman Bansal, Rounak Saha, Manav Nitin Kapadnis, Millon Madhur Das, Pawan Goyal, Niloy Ganguly
AI summary
Overview
- Research area: Multimodal natural language processing / vision-language modeling, specifically computational humor (satire, sarcasm, memes) and prompting-based reasoning with large vision-language models (VLMs).
- Technical level: Intermediate. The core idea is conceptually simple (turn a cause-effect graph into code inside a prompt), but the paper assumes familiarity with prompting baselines (Chain-of-Thought, Chain-of-Draft, scene graphs), causal graph formalism, and information-theoretic comparison measures.
- Scope: The paper proposes a VLM-agnostic prompting framework, CaRGo-T, that makes a VLM first emit an explicit causal reasoning graph as code and then answer a humor question conditioned on that graph, and evaluates it against reasoning-based baselines on humor understanding and humor detection tasks (arXiv:2608.23172v2, CC BY 4.0).
What This Paper Is About
Current vision-language models handle many multimodal tasks well but still struggle with humor, because jokes depend on non-obvious interactions among people, objects, abstract concepts and events — exactly the kind of relational, cause-and-effect structure that flat natural-language rationales tend to flatten into a single literal reading. The paper's goal is to force a VLM to make those causal relationships explicit, as a small code-based graph, before it produces an answer, and to test whether that improves both explaining why something is funny and deciding whether it is funny at all.
Key Contributions
- CaRGo-T, a VLM-agnostic framework for constructing and using explicit causal reasoning graphs (CRGs) as the reasoning component in visual humor tasks. A CRG is described as a restricted subclass of causal graphs retaining only cause–effect links plus lightweight metadata about objects, concepts, events and participants, with no probabilistic parameters.
- A code-based serialization of the graph. The VLM writes the CRG as a piece of code in the prompt, and the same or a different VLM then interprets that code to produce the final answer, in either a zero-shot or in-context-learning setting. The paper states it is the first work in multimodal humor comprehension to explicitly model causal graphs or cause–effect relationships for this purpose.
- Comprehensive experiments across humor understanding (Satirical Image Understanding on the YesBut Dataset, Meme Caption Generation on the MemeCap Dataset) and humor detection (Satirical Image Detection on YesBut, Multimodal Sarcasm Detection on MMSD 2.0), using GPT-4o, GPT-4o-mini and the open-source MiniCPM-V-2_6 model.
- An analysis of the generated reasoning component, comparing how much new lexical and semantic information CaRGo-T's reasoning contains relative to baselines, and how often the ground-truth answer can be logically inferred from it.
Main Findings
- Zero-shot gains on humor understanding. CaRGo-T improved the average score by 0.72% on the Satire task and 5.81% on the Meme task compared to the best baseline when using MiniCPM. Highlighted average scores for CaRGo-T were 0.3504 (MiniCPM, Satire), 0.3295 (MiniCPM, Meme), 0.3632 and 0.3318 (GPT-4o-mini), and 0.3726 and 0.3321 (GPT-4o).
- Consistent improvement for GPT-4o-mini. The paper reports CaRGo-T beating all baselines across the board with GPT-4o-mini, unlike MiniCPM, attributing this to better code generation and a larger context window (GPT-4o-mini: 128k tokens; MiniCPM: 32k tokens) despite similar model sizes.
- In-context learning helps with diminishing returns. With GPT-4o, CaRGo-T improved over CoT by 11.66% in the 0-shot setting, 10.14% in 2-shot and 5.86% in 5-shot. The paper notes that more in-context examples do not necessarily improve a VLM's understanding, especially as model size grows.
- Better gains on Meme Captioning than Satire Understanding, which the authors suggest is because the meme title provides supporting text that gives better supervision for constructing the causal reasoning graph.
- Humor detection improves more modestly. On MMSD 2.0 sarcasm detection with GPT-4o, CaRGo-T reached 49.48% accuracy and 62.20% F1 in 0-shot (2.93% and 0.83% over the best baseline), 49.88% and 62.38% in 2-shot, and 49.91% and 62.41% in 6-shot. On YesBut satire detection, CaRGo-T reached 43.18% accuracy and 59.97% F1 in 0-shot, 44.63% and 61.01% in 2-shot, and 45.57% and 61.56% in 6-shot. Sarcasm detection gains exceed satire detection gains, which the authors again attribute to the supporting text in the sarcasm task.
- The reasoning component carries more information. KL divergence between token distributions was 0.21 from CaRGo-T's reasoning to CoT's versus 0.19 in the reverse direction; 0.21 versus 0.20 against CoD; and 0.25 versus 0.25 against CCoT. The sentence-similarity dissimilarity score was 1 against CoT (0.85 reverse), 1 against CoD (0.84 reverse), and 0.98 against CCoT (0.94 reverse).
- Ground truth is more often inferable from CaRGo-T's reasoning. Using an LLM-as-a-judge setup with GPT-4, the InferScore was 45.11 for CaRGo-T versus 40.78 for CoT, 40.68 for CoD and 37.64 for CCoT.
- Manual rectification matters. Ablations on GPT-4o showed CaRGo-T beating an unrectified variant of itself at both 2-shot (average 0.3911 versus 0.3771) and 5-shot (0.3901 versus 0.3758), and beating a variant that adds a detailed definition of the causal reasoning graph to the prompt (0.3726 versus 0.3687 in 0-shot; 0.3911 versus 0.3844 in 2-shot; 0.3901 versus 0.3877 in 5-shot). The authors suggest including the graph definition alongside the task query may confuse the VLM.
- Headline improvement range. The abstract reports roughly 1–20% improvement in Humor Understanding and roughly 1–3% in Humor Detection over other reasoning-based baselines, with improvements reported as significant under an independent two-sample t-test.
Methodology in Plain English
The framework starts from an ordinary task prompt — for example, "Why is this image funny/satirical?" The prompt is then extended with an instruction to first build a causal reasoning graph linking the objects, people and entities in the image (and any input text) as a piece of code, and only then to give the final answer. The expected output format is Code: followed by the graph, then Final Answer:.
For the zero-shot setting, nothing else is added, and the method depends on the VLM's own code-generation ability, which is why the authors mostly use closed proprietary models (GPT-4o, GPT-4o-mini) there. For the in-context-learning setting, a few examples from outside the test set are included, each consisting of the input, a causal reasoning graph, and the ground-truth answer. The graphs in those examples are first drafted by GPT-4o conditioned on the input and the known correct answer, then manually corrected into a fixed structure: identify entities (object, person, abstract concept or event), list each entity's properties, then list cause-effect pairs. The comparison in the appendix shows how GPT-4o's draft graph was restructured into the rectified form.
Baselines are all VLM-agnostic: Vanilla prompting, Chain-of-Thought, Chain-of-Draft (concise draft summaries, zero-shot only), and Compositional Chain-of-Thought (which injects an automatically generated scene graph). Humor understanding is scored with ROUGE-L, BLEU and BERTScore plus their mean (with the YesBut satire metrics averaged across the dataset's 3 stages), and humor detection with accuracy and macro-F1.
To examine why the method works, the authors compare reasoning text in three ways: token-level KL divergence (tokens produced by whitespace splitting and lowercasing), a sentence-similarity-based dissimilarity score using Sentence-BERT embeddings with a cosine similarity threshold of 0.5, and an InferScore where GPT-4 judges whether the ground-truth answer can be logically inferred from the reasoning. The abstract frames this as an analysis of the "mutual information" of the reasoning component, while the reported measures are the KL divergence, the dissimilarity score and the InferScore.
Test data used: 1,079 satirical images from YesBut for satire understanding (with 5 holdout examples for in-context selection), 2,541 images from YesBut for satire detection (1,081 satirical, 1,460 non-satirical, 6 held out), 559 MemeCap test samples, and 2,409 MMSD 2.0 test samples (1,037 sarcastic, 1,372 non-sarcastic). Open-source experiments ran on 2 NVIDIA L40 GPUs with 48GB VRAM each. Code is released at https://github.com/abhi1nandy2/CaRGo-T.
Why This Matters
Research impact. The paper argues that natural-language rationales (Chain-of-Thought) collapse in subjective, emotion-laden contexts and that self-reflection-style multimodal extensions produce noisy, misaligned rationales. Replacing free-form rationale text with an explicit, inspectable cause-effect structure gives a different kind of reasoning scaffold, and the reasoning-component analysis offers a template for measuring whether a reasoning trace actually contains task-relevant information rather than just being longer.
Real-world applications.
- Content moderation and platform safety systems that need to judge sarcastic or satirical posts from image-plus-text pairs, where literal readings are misleading.
- Accessibility tools that explain why a meme, cartoon or satirical image is intended to be funny for users who miss the cultural or social cue.
- Advertising and brand-safety tooling that must flag sarcastic or satirical user content about a product.
- Assistive communication and social-skills training tools that need to surface the implicit relational dynamics behind an image.
Industry relevance. The framework is model-agnostic and requires no fine-tuning — only prompt engineering plus a small number of curated in-context examples — which makes it directly deployable on top of whichever commercial or open VLM an organization already uses. The finding that a small open-source model gains (MiniCPM averaged 0.3504 versus 0.3479 for the best baseline on Satire, and 0.3295 versus 0.3114 on Memes) and that GPT-4o-mini benefits consistently matters for cost-sensitive deployments, though the zero-shot variant leans on strong code-generation ability, favoring proprietary models.
Future Directions
- Remove the manual step. In-context examples currently require humans to rectify GPT-4o-drafted causal reasoning graphs; automating that rectification reliably would make the framework cheaper to apply to new tasks.
- Explain why more in-context examples stop helping. The paper observes diminishing returns as examples increase and as model size grows, but the underlying cause is not established.
- Improve humor detection headroom. Detection gains are small in absolute terms (highest reported accuracy is 49.91% on MMSD 2.0 and 45.57% on YesBut satire detection), leaving substantial room before these tasks are solved.
- Test whether causal reasoning graphs transfer beyond humor, and whether they help with other subjective or affect-laden multimodal tasks — the paper notes no prior work has used cause-effect relations to boost open-ended multimodal reasoning more broadly.
Target Audience
Researchers and practitioners working on multimodal reasoning, computational humor, sarcasm and satire detection, and prompting strategies for VLMs. It is also useful for engineers who want a no-training, prompt-level intervention to improve VLM behavior on subjective multimodal judgments, and for readers interested in how to measure whether a model's reasoning trace is actually informative through information-theoretic and LLM-as-a-judge comparisons.
Authors’ abstract
Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal tasks. However, understanding humor remains challenging because humorous content often depends on subtle interactions among entities, events, context, and implicit relationships across image and text modalities. These interactions can involve complex chains of reasoning that are difficult to capture through conventional prompting or linear chain-of-thought reasoning. In this work, we propose CaRGo-T (Causal Reasoning Graph-of-Thought), a reasoning framework that represents the causal and contextual relationships underlying multimodal humor as a lightweight graph-based reasoning structure. The graph is serialized into a code-based representation generated by a VLM, which can subsequently be interpreted by the same or a different VLM to produce the final prediction in zero-shot or in-context learning settings. We evaluate CaRGo-T on humor understanding and humor detection across four datasets spanning diverse forms of comedic content, including satire, sarcasm, and memes. Experiments with state-of-the-art commercial and open-source VLMs show that CaRGo-T consistently improves performance over existing reasoning-based baselines, achieving gains of approximately 1-20% on humor understanding and 1-3% on humor detection. Further analysis using mutual information indicates that the reasoning representations produced by CaRGo-T contain more information relevant to the target output than those generated by baseline reasoning approaches. Code is available at https://github.com/abhi1nandy2/CaRGo-T.