Research
Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances
Overview Research area: Natural Language Processing, specifically multimodal (text + image) complaint understanding in multi-turn customer-support dialogues. Technical level: Advanced. The paper assum

- arXiv
- 2511.14693
- Published
- 2025-11-18
- Authors
- Rishu Kumar Singh, Navneet Shreya, Sarmistha Das, Apoorva Singh, Sriparna Saha
AI summary
Overview
Research area: Natural Language Processing, specifically multimodal (text + image) complaint understanding in multi-turn customer-support dialogues.
Technical level: Advanced. The paper assumes familiarity with transformer encoders (BERT, ViT), cross-modal attention, Mixture-of-Experts (MoE) routing, Chain-of-Thought prompting, and multi-objective loss design.
Scope: The paper introduces CIViL, a benchmark dataset of annotated multimodal customer-support conversations, and VALOR, a two-phase validation-aware MoE framework for jointly predicting complaint aspect categories and severity levels.
What This Paper Is About
Most prior complaint analysis works on short, single-turn, text-only inputs such as tweets or product reviews, even though real customers now routinely attach visual evidence (screenshots, photos of damaged items) and elaborate their issues across multiple turns of a support conversation. This paper reframes complaint analysis as a fine-grained multimodal classification task over multi-turn dialogues, where the system must identify both the aspect category (e.g., Software, Hardware, Packaging) and the severity level (e.g., No Explicit Reproach, Disapproval, Blame, Accusation). The goal is to build a framework that fuses conversational text with aligned images and validates its own expert predictions before producing a final answer.
Key Contributions
- Task formulation: The authors define and investigate fine-grained multimodal complaint understanding in multi-turn dialogues, targeting aspect category detection (ACD) and severity detection (SD) jointly.
- CIViL dataset: They introduce CIViL (Customer Interactions with Visual and Linguistic signals), a benchmark of 2,004 customer-support conversations (7,101 utterances, 4,478 images) built by extending a subset of the Kaggle Customer Support on Twitter corpus with fine-grained aspect and severity annotations and enriched with topically aligned images.
- VALOR framework: They propose Validation-Aware Learner with Expert Routing, a two-phase Mixture-of-Experts architecture combining cross-modal fusion, a learnable Semantic Alignment Score, Chain-of-Thought expert reasoning, and a secondary validation MoE with meta-fusion.
- Benchmarking: They report that VALOR outperforms thirteen baseline multimodal models (zero-shot, few-shot, and fully fine-tuned paradigms) on the CIViL dataset across Accuracy and macro F1 for both aspect and severity tasks, supported by ablation, human evaluation, and qualitative analysis.
Main Findings
- Dataset composition: CIViL contains 2,004 conversations, 7,101 utterances, 4,478 images, 3,825 customer utterances, and 3,276 support-agent utterances. Average user utterances per conversation is 2.74 and average support-agent utterances is 1.49.
- Label distributions: Severity levels are Blame (799), Disapproval (486), Accusation (484), and No Explicit Reproach (235). Aspect categories are Software (1,662), Quality (117), Hardware (112), Service (77), Price (23), and Packaging (13).
- Annotation agreement: Fleiss' Kappa was 0.68 for aspect categories and 0.75 for severity levels, described as substantial consistency. Three annotators (one Ph.D. researcher and two postgraduate scholars) labeled each dialogue independently, with disputes resolved in collaborative review sessions.
- Strongest baseline: Gemma-3 (9B) was the best-performing baseline, reaching 0.69 ACD accuracy and 0.66 ACD F1, with 0.65 SD accuracy and 0.66 SD F1. The weakest baselines listed were ViLT and VisualBERT.
- VALOR's headline results: VALOR achieved 81.94% aspect accuracy, 76.96% aspect F1, 72.51% severity accuracy, and 67.91% severity F1. The paper reports this as absolute improvements of 12.94% (aspect) and 6.51% (severity) over Gemma-3.
- Validation module matters most: Adding the Validation MoE raised aspect accuracy from 73.74% to 81.94%, a gain of 8.2%. The full VALOR configuration used CoT experts, validation enabled, learnable SAS, and Top-2 routing.
- Learnable alignment beats static alignment: The learnable Semantic Alignment Score outperformed cosine similarity and alignment-agnostic ("none") settings in the ablation table.
- Chain-of-Thought experts win: CoT experts outperformed both Transformer-based and MLP-based experts. MLP experts produced the weakest results; Transformer experts were competitive but the authors note they fall short in interpretability and sequential inference.
- Statistical significance: The authors report a Student's t-test with the null hypothesis rejected at p < 0.05.
- Human evaluation: Over 200 randomly selected test samples using a win-loss-draw protocol against Gemma-3, DeepSeek-VL, and Flash Gemini, VALOR achieved the highest win rates at 42.3% for aspect identification and 38.5% for severity classification, with the lowest loss rates of 18.7% and 22.1%, respectively.
- Error analysis: Two key failure modes are reported — subjective severity interpretation (variability in user tone can cause underestimation or misclassification) and class imbalance (heavy representation of "software" versus sparse categories like "price").
- Qualitative strength: VALOR handles multi-aspect complaints (e.g., a "battery issue and slow software" complaint resembling hardware–disapproval and software–accusation), and the validation layer intervenes in low-confidence primary-expert cases.
Methodology in Plain English
The researchers built a dataset first. They started from the publicly available Kaggle Customer Support on Twitter corpus, focused on Apple Support conversations (which make up 14% of that data), and filtered for two-speaker dialogues between 2 and 10 utterances. From that pool they randomly sampled 2,004 conversations and had three annotators label each with aspect categories and severity levels, using written guidelines and a reference set of 50 sample conversations.
For the images, they scraped 4,478 pictures from X (using Scrapy) and Reddit (using PRAW), targeting 15 complaint themes across nine subreddits. Filters required a general keyword such as "iphone," "apple," or "ios," a minimum upvote score, a direct image URL, and resolution above 50,000 total pixels. Then a CLIP-based semantic matching algorithm assigned the most topically aligned image to each conversation, retaining only high-confidence image–conversation pairs. The appendix also describes BLIP being used for analysis alongside CLIP for matching.
The model, VALOR, works in two phases. In phase one, text is tokenized with BERT-base-uncased (vocabulary 30,522) and truncated to 512 tokens, while images are resized to 224×224 and passed through a ViT-patch16 module producing 196 patches. A 12-layer, 12-head BERT (hidden size 768) and a ViT-base encoder produce the text and image representations, which are fused with 8-head cross-modal attention and mean-pooled into one multimodal embedding. Separately, a Semantic Alignment Score is computed by projecting the text and image [CLS] vectors into a shared 512-dimensional space through two-layer MLPs with GELU activation, then passing the concatenated, layer-normalized result through another MLP with tanh activation to yield a scalar in [-1, 1]. The fused embedding is routed by a learned gating function to 4 Chain-of-Thought experts built on DeepSeek-6.7B (hidden size 4096), sampled with temperature 0.5, top-k 30, and top-p 0.9. A load-balancing penalty discourages routing collapse.
In phase two, 2 validation experts (DeepSeek transformer stacks with 32 layers and hidden dimension 4096) re-examine the joint embedding. The authors use a three-part metric system — alignment (cosine similarity between expert logits, averaged as R_avg), dominance (correlation between primary MoE and validation logits), and complementarity (entropy over softmax-normalized validation logits, averaged as U_avg). A meta-fusion network concatenates the primary logits, validation logits, the alignment score, routing entropy, R_avg, dominance, and U_avg, then passes them through a 3-layer MLP with hidden sizes (768, 384, C_a), ReLU activations, and dropout 0.1. The final logits get a small adjustment from the alignment score (λ_s = 0.1).
Training combines label-smoothed cross-entropy for aspect and severity (ε_ls = 0.15) with the validation loss, a semantic alignment margin loss (μ = 0.3), the load-balancing loss, and regularizers on alignment (τ_R = 0.3), dominance (τ_S = 0.5), and complementarity (τ_U = 1.5). Experiments used a 70/10/20 train/validation/test split, an NVIDIA RTX 3090 GPU, AdamW with learning rate 2 × 10⁻⁵, warm-up and cosine decay, batch size 16, dropout 0.5, gradient clipping max norm 1.0, up to 20 epochs with early stopping (patience 5), and random seed 42. Performance was measured by Accuracy and macro F1 for both tasks.
Why This Matters
Research impact: The work shifts complaint analysis away from isolated, text-only short posts toward multimodal, multi-turn dialogue, and it contributes both a benchmark dataset and a modular architecture that other researchers can ablate and extend. The validation-aware MoE design — where a second set of experts audits the first — offers a reusable pattern for other fine-grained multimodal classification problems.
Real-world applications:
- Intent-based ticket routing: Structured aspect predictions can automatically send complaints to the right support queue (e.g., hardware vs. software vs. packaging).
- Escalation prediction: Severity levels such as Blame and Accusation can flag conversations that need urgent human attention.
- Service analytics: Aggregated aspect–severity pairs give companies a quantified view of which product dimensions generate the most intense dissatisfaction.
- Product feedback loops: Visual evidence attached by users (broken screens, battery drain screenshots) can be systematically linked back to design and quality teams.
Industry relevance: The paper aligns its motivation with UN SDG 9 (Industry, Innovation and Infrastructure) for scalable, context-aware service infrastructure and SDG 12 (Responsible Consumption and Production) for more responsive product design and accountability in consumer services. Its authors state the work is conducted solely for the research community and is not intended for commercial use, and that neither the authors nor annotators intend to defame any company.
Future Directions
- Multilingual support: The authors explicitly list extending the framework to multilingual scenarios as future work.
- Additional multimodal signals: They plan to incorporate further signal types beyond the text-plus-image setting studied here.
- Speaker roles and temporal dependencies: Modeling who is speaking and how complaints evolve across dialogue turns is identified as a route to better applicability across service contexts and user populations.
- Class imbalance and subjective severity: The error analysis highlights sparse aspect categories (e.g., Price with 23 instances, Packaging with 13) and ambiguous severity wording as open problems that a more balanced or more nuanced severity model could address.
Target Audience
This paper is most useful to NLP and multimodal machine learning researchers working on complaint analysis, dialogue systems, or Mixture-of-Experts architectures; to practitioners in customer support automation and service analytics who need structured, fine-grained outputs from conversations; and to dataset builders interested in annotation protocols for multimodal, multi-turn corpora. Readers without a background in transformer architectures or MoE routing will find the methodology sections dense, though the introduction, dataset description, results, and error analysis remain accessible.
Authors’ abstract
Existing approaches to complaint analysis largely rely on unimodal, short-form content such as tweets or product reviews. This work advances the field by leveraging multimodal, multi-turn customer support dialogues, where users often share both textual complaints and visual evidence (e.g., screenshots, product photos) to enable fine-grained classification of complaint aspects and severity. We introduce VALOR, a Validation-Aware Learner with Expert Routing, tailored for this multimodal setting. It employs a multi-expert reasoning setup using large-scale generative models with Chain-of-Thought (CoT) prompting for nuanced decision-making. To ensure coherence between modalities, a semantic alignment score is computed and integrated into the final classification through a meta-fusion strategy. In alignment with the United Nations Sustainable Development Goals (UN SDGs), the proposed framework supports SDG 9 (Industry, Innovation and Infrastructure) by advancing AI-driven tools for robust, scalable, and context-aware service infrastructure. Further, by enabling structured analysis of complaint narratives and visual context, it contributes to SDG 12 (Responsible Consumption and Production) by promoting more responsive product design and improved accountability in consumer services. We evaluate VALOR on a curated multimodal complaint dataset annotated with fine-grained aspect and severity labels, showing that it consistently outperforms baseline models, especially in complex complaint scenarios where information is distributed across text and images. This study underscores the value of multimodal interaction and expert validation in practical complaint understanding systems. Resources related to data and codes are available here: https://github.com/sarmistha-D/VALOR