Research
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation Overview Research area: Computer vision / affective computing — emotion-aware video understanding and
- arXiv
- 2511.11002
- Published
- 2025-11-14
- Authors
- Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang
AI summary
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and GenerationOverview
- Research area: Computer vision / affective computing — emotion-aware video understanding and video generation, with a focus on stylized and non-realistic content.
- Technical level: Intermediate. The dataset construction and analysis are accessible to a general reader, but the generation benchmarks assume familiarity with text-to-video (T2V) and image-to-video (I2V) models, diffusion fine-tuning, and LoRA.
- Scope: The paper builds and releases EmoVid, a large-scale, multimodal, emotion-labeled video dataset spanning animation, movie, and sticker content, then uses it as a benchmark and as fine-tuning data to make video generation models more emotionally expressive.
What This Paper Is About
Video generation models have become good at visual coherence and motion, but they largely ignore emotion, and the datasets needed to study video emotion are mostly small, single-modality, or restricted to realistic human faces and dialogue. The authors create EmoVid, described as the first multimodal emotion-annotated video dataset for artistic media, covering cartoon animations, movie clips, and animated stickers, with emotion labels, visual attributes, and text captions. They then use it to benchmark existing T2V and I2V models on emotional accuracy and to fine-tune the Wan2.1 model so that generated videos better match a requested emotion.
Key Contributions
- The EmoVid dataset. A large-scale, emotion-labeled video dataset focused on stylized and non-realistic content, containing 22,758 clips across three types (Animation, Movie, Sticker) with a total duration of 140,580 seconds, annotated with discrete emotion labels, low-level visual attributes (brightness, colorfulness, hue), audio tracks (for Animation and Movie), and text captions.
- A benchmark, evaluation metrics, and protocol. A T2V and I2V benchmark of 240 videos (10 representative videos per emotion category across 3 video types and 8 emotions), evaluated with FVD, CLIP score, SD score, temporal flicker, and emotion accuracy (EA-2cls and EA-8cls).
- Spatial and temporal analysis of emotion in video. Studies of emotion distributions, color–emotion correlations, the temporal transition structure of emotion in consecutive movie clips, and the relationship between captions and emotion labels.
- Emotion-conditioned generation via fine-tuning. LoRA fine-tuning of the Wan2.1 T2V and I2V models on EmoVid, improving emotional accuracy and emotional expressiveness, plus a pipeline for generating animated stickers of a character expressing a chosen emotion and multi-LoRA combinations with style/identity LoRAs.
Main Findings
- Dataset composition: EmoVid totals 22,758 clips and 140,580 seconds. It breaks down into 2,807 animation clips (average 5.12 s, SD 2.65), 13,255 movie clips (average 8.75 s, SD 4.71), and 6,696 animated stickers (average 2.91 s, SD 2.18). Mean clip duration is 6.18 s (SD 4.53), with the shortest clip 0.18 s and the longest 29.98 s. Movies make up 58.24% of the data, stickers 29.42%, and animation 12.33%.
- Human vs. machine labels: 10,049 clips carry human-annotated emotion labels (282 animation clips, 2,771 movie clips, and 6,996 sticker clips); the rest are labeled by a fine-tuned VLM. On a 1% validation set, the overall Fleiss' kappa was 0.3704, the average inter-human Cohen's kappa was 0.311, and the average human–VLM kappa was 0.301 — a small difference (under 4%) between human–human and human–VLM agreement.
- VLM labeler selection: On EmoSet, fine-tuned NVILA-Lite-2B reached 87.5% accuracy, versus 66.68% for ResNet50, 45.46% for VGG-16, 26.88% for TinyLLaVA-Phi-2-SigLIP-3.1B, 38.75% for the fine-tuned TinyLLaVA variant, and 57.5% for non-fine-tuned NVILA.
- Emotion distribution is imbalanced in animation and movie content: Anger and sadness appear frequently, while amusement and awe are underrepresented in animation and movie types — which the authors attribute to reflecting real-world emotional distributions in the collected content. Per-category totals across all styles range from 1,768 (Awe) to 4,756 (Anger).
- Color relates to emotion: The positive-to-total emotion ratio trends upward with colorfulness and brightness. Positive-valence categories are brighter and slightly more colorful than negative-valence ones, and high-arousal emotions tend to be darker but more colorful than low-arousal emotions. ANOVA indicates the effects are statistically significant (p < 0.01) but small (η² < 1%), so color is useful for stylistic guidance or weak supervision, not standalone classification.
- Emotions are temporally persistent with intra-valence drift: A first-order Markov transition matrix from consecutive movie clips shows strong self-persistence — notably fear (0.53), anger (0.46), and amusement (0.46). Transitions within the same valence polarity are more frequent (typically 0.08–0.18) than across polarity (<0.08). Negative emotions show a chain-like escalation, such as sadness → fear/anger and fear → anger, which the authors describe as a possible "defense-attack" progression. Overall they summarize the trajectory as "hold, intra-valence drift, arousal leap."
- Captions align with emotion: Vocabulary analysis of captions using 2–4 word n-grams shows affectively consistent phrasing — for example, "Merry Christmas" and "Peace written" for positive categories, and phrases like "Man holding gun aiming" and "Boxing glove" for negative categories. Caption sentiment polarity corresponds to emotion label polarity.
- Fine-tuning improves emotional accuracy: For T2V, WanVideo (after) improves EA-2cls from 84.17 to 88.33 and EA-8cls from 44.16 to 48.33 over WanVideo (before), with FVD improving from 594.3 to 573.7 and CLIP from 0.2982 to 0.3021, though flicker rises from 0.0091 to 0.0143. For I2V, WanVideo (after) improves EA-2cls from 91.25 to 94.58 and EA-8cls from 71.30 to 76.25, versus DynamiCrafter512 (90.41 / 71.25), HunyuanVideo (89.17 / 70.00), and CogVideoX (90.83 / 70.83). The paper's text also cites an I2V emotion classification accuracy of 92.08% (2-class) and 72.92% (8-class), figures that differ from the values in its results table.
- Qualitative gains: The baseline Wan2.1 I2V model often produced neutral or mismatched expressions, while the fine-tuned model showed more precise emotional articulation, including heightened facial expressions, contextual cues, and mood-consistent motion. The fine-tuned model also generates character stickers for different emotions, and combining the emotion LoRA with identity/style LoRAs produces videos with specified emotional attributes (for example, Studio Ghibli style).
Methodology in Plain English
The authors first assembled video data from three sources: cartoon face clips adapted from the MagicAnime dataset, movie clips retrieved using metadata and code from Condensed Movies (segmented with PySceneDetect and retained only if between 4 and 30 seconds), and animated stickers searched through the Tenor API using eight emotion labels and their synonyms, then manually verified. Stickers were filtered down from 9,633 initial clips to 6,696.
For emotion labels they adopted the Mikels eight-emotion scheme (amusement, awe, contentment, excitement, anger, disgust, fear, sadness) and placed the categories on the valence-arousal model. Because human labeling is expensive, they used a human–machine hybrid: 12 annotators were recruited, with two validating the sticker labels and ten labeling 20% of the animation and movie data (3,000 movie clips and 600 animation clips), each clip labeled independently by three annotators with audio, kept only if at least two agreed. They first compared labelers on EmoSet and found that a fine-tuned NVILA-Lite-2B performed close to human level, so the remaining 80% of animation and movie labels were produced by that model; captioning used NVILA-8B-Video. They also computed per-clip brightness, colorfulness, and hue from the HSV color space by sampling every 20 frames, using circular statistics for hue.
They then analyzed the dataset: a t-SNE visualization of video features (animation and movie clusters separate, stickers overlapping both), emotion distributions by domain, a Markov transition matrix over consecutive movie clips, color–emotion statistics per category, and caption polarity and keyword analysis with NLTK and CountVectorizer. For evaluation, they built a benchmark of 240 videos and tested four T2V models (VideoCrafter-V2, HunyuanVideo, CogVideoX-5B, Wan2.1-T2V-14B) and four I2V models (DynamiCrafter512, HunyuanVideo-I2V, CogVideoX-I2V, Wan2.1-I2V-480P) on their own prompts with an emotion label appended. Finally, they fine-tuned both Wan2.1 variants with LoRA using the DiffSynth Studio framework on an H20 GPU with 96 GB memory, rank 32, learning rate 1e-4, 3 epochs, batch size 1, training on 2,727 animation clips, 8,000 movie clips, and 6,616 sticker clips.
Why This Matters
EmoVid addresses a documented gap: existing emotion video datasets are small, missing modalities, or restricted to realistic faces and dialogue, which limits their transfer to creative, stylized video generation. By pairing a dataset with a benchmark and a working fine-tuning recipe, the paper gives the affective video computing community a shared evaluation target where emotional accuracy is treated as the primary criterion rather than realism or temporal fidelity. The dataset is released for non-commercial research use by academic institutions only, with redistribution and commercial use prohibited.
Real-world applications the paper points to:
- Generating emotion-conditioned animated stickers or memes of any character, aimed at social media communication.
- Supporting animation and film production workflows where emotional clarity drives narrative.
- Emotion-aware avatar generation and expressive media content synthesis.
- Controllable video editing driven by emotional cues, and emotion-aware video pacing and editing strategies informed by the transition analysis.
Industry relevance: the work targets creative and entertainment pipelines where emotional expressiveness, not photorealism, is the product — animation studios, social platforms built around stickers and short-form expressive content, and teams integrating emotion control into generative video tooling.
Future Directions
- Moving beyond one emotion per clip. The authors note that their work assumes each clip conveys a specific emotion, while real-world expressions can be highly detailed and composite, so finer-grained or mixed-emotion modeling remains open.
- Better use of audio. The audio component is only partially leveraged; the authors state that building a truly unified video–audio–text multimodal model is future work.
- Extending the temporal and cross-domain analysis. The transition analysis relies on consecutive movie clips; whether the same "hold, intra-valence drift, arousal leap" trajectory holds for animation and stickers is not established in the reported content.
- Broadening evaluation. Whether color attributes or caption semantics can serve as stronger supervision signals is left open, given that ANOVA effect sizes for color were small (η² < 1%), and the discrepancies between the reported I2V emotion accuracy figures in the text and the results table are not explained.
Target Audience
Researchers and practitioners in affective computing, multimodal learning, and generative video. It is most useful for those building or evaluating emotion-aware video generation and editing systems, and for teams working with stylized, animated, or cinematic content where emotional expressiveness matters more than realism. Dataset and benchmark builders will find the human–machine labeling protocol and the VLM labeler comparison directly applicable, while readers primarily interested in realistic facial emotion recognition will find it relevant mainly as a contrasting, stylized resource.
Authors’ abstract
Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to bridge emotion understanding with generative tasks, particularly for stylized and non-realistic contexts. To address this gap, we introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for creative media, which includes cartoon animations, movie clips, and animated stickers. Each video is annotated with emotion labels, visual attributes (brightness, colorfulness, hue), and text captions. Through systematic analysis, we uncover spatial and temporal patterns linking visual features to emotional perceptions across diverse video forms. Building on these insights, we develop an emotion-conditioned video generation technique by fine-tuning the Wan2.1 model. The results show a significant improvement in both quantitative metrics and the visual quality of generated videos for text-to-video and image-to-video tasks. EmoVid establishes a new benchmark for affective video computing. Our work not only offers valuable insights into visual emotion analysis in artistically styled videos, but also provides practical methods for enhancing emotional expression in video generation.