Skip to content
AI.info

Research

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

VideoNorms: Benchmarking Socio-Cultural Norm Understanding of Video Language Models Overview Research area: Video-language model evaluation, cultural competence, multimodal social reasoning, benchmark

arXiv
2510.08543
Published
2025-10-09
Authors
Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan

AI summary

VideoNorms: Benchmarking Socio-Cultural Norm Understanding of Video Language Models

Overview

Research area: Video-language model evaluation, cultural competence, multimodal social reasoning, benchmark construction (computer vision / NLP).

Technical level: Intermediate. Readers should be comfortable with video large language models (VideoLLMs), binary classification metrics such as Macro F1, inter-annotator agreement statistics, and the general logic of hierarchical regression models.

Scope: The paper introduces a human-AI annotation framework, a dataset of over 3,000 socio-cultural norm judgments drawn from US and Chinese television, a filtered 542-instance benchmark, and a cross-cultural evaluation of seven open-weight VideoLLMs on norm adherence/violation prediction and evidence generation.

What This Paper Is About

VideoLLMs are deployed globally, but their ability to understand culturally specific social norms has received less attention than object recognition, temporal reasoning, or narrative understanding. The authors build a benchmark for judging whether a video language model can tell when characters in a clip adhere to or violate a socio-cultural norm, using comparable US and Chinese TV shows. The goal is to measure and characterize cross-cultural performance gaps and to test whether the models can point to the verbal and non-verbal evidence that justifies their judgment.

Key Contributions

  1. A human-AI collaboration framework for video norm annotation. A large VideoLLM (Gemini-2.0-Pro) generates candidate norms, norm categories, adherence/violation labels, and evidence from 15-second clips; three trained monocultural annotators per culture then validate or edit every field. Annotators could only advance after confirming agreement with the candidate annotation when no edits were made.

  2. The VideoNorms dataset. Over 3,000 human annotations of socio-cultural norms drawn from eight shows — four US and four Chinese — spanning formal (workplace) and informal settings across drama and comedy genres. The US portion contains 266 video clips, 511 annotated items, and 1,533 total annotations; the Chinese portion contains 249 clips, 500 items, and 1,500 annotations.

  3. VideoNorms-Benchmark. A high-reliability subset of 542 instances (273 US, 269 Chinese) where all three annotators agreed on context, category, subject, specific norm, and adherence/violation; for 2-to-1 splits, two additional meta-annotators were recruited and a 4/5 majority vote applied.

  4. An empirical study of seven open-weight VideoLLMs on two tasks — binary adherence/violation prediction and explainable prediction (verbal + non-verbal evidence) — analyzed with hierarchical linear modeling, plus ablations over input modality and model size.

Main Findings

  • Models show a cultural performance disparity in adherence prediction. The hierarchical linear model estimates that, across VideoLLMs and shows, the odds of a correct prediction for norm adherence identification are 57% higher for US than for China (OR = 1.57, p = 0.006). No equivalent effect was identified for violation prediction, which the authors attribute possibly to the lower sample size of violations.

  • Non-verbal evidence is harder than verbal evidence. The hierarchical linear regression for evidence scores shows models obtain significantly lower scores for non-verbal evidence than verbal evidence, suggesting the language backbone handles transcripts better than the multimodal components handle visual social cues.

  • Human verification revealed a large US/China gap in candidate annotation quality. Combined annotator change percentages ranged from 18.28% to 26.04% for US shows but 42.35% to 53.39% for Chinese shows, which the authors read as a caution against fully automated cultural norm extraction, especially for cultures under-represented in training data. Per-field edits were especially high for Chinese: verbalEvidence 64.54%, specificNorm 57.88%, context 56.82%, nonverbalEvidence 53.89%, normActors 53.83%, normCategory 50.83%, normAdherence 26.28%, timestampEnd 11.78%, timestampStart 4.86% (US figures were 23.35%, 16.99%, 4.99%, 19.71%, 29.44%, 14.59%, 13.62%, 3.05%, 2.40% respectively).

  • Agreement varied by show. Free-marginal multirater Kappa for adherence/violation was highest for workplace dramas (Suits κ_free = 0.89, Best Partner κ_free = 0.95). Fleiss's kappa and change percentages by show: Friends 0.61 / 23.36%, Big Bang Theory 0.71 / 18.28%, The Office 0.59 / 26.04%, Suits 0.76 / 20.62%, iPartment 0.70 / 50.00%, Home with Kids 0.59 / 46.89%, Amazing Night 0.73 / 53.39%, Best Partner 0.66 / 42.35%.

  • Agreement statistics differ sharply across cultures. In the US data, all three annotators agreed on 54.0% of items, 2 agreed and 1 differed on 37.8%, and all three disagreed on 8.2%. In the Chinese data these figures were 14.2%, 45.0%, and 40.8%. The benchmark retains class imbalance: US adherence 63.7% / violation 36.3%; Chinese adherence 82.9% / violation 17.1%.

  • No single model leads everywhere. Intern3.5-VL slightly outperforms on US classification (Task 1 Macro F1 71.5 ±5.4). Llava-OneVision excels on Chinese classification. For verbal evidence on correct predictions, Intern3-VL achieves the highest scores, which the authors attribute to fine-grained visual grounding. VideoChatR1 preserves classification performance while generating explanations. Intern3.5-VL drops from Macro F1 71.5 (Task 1) to 67.8 (Task 2) on US norms when evidence generation is required.

  • Video modality is necessary. For Intern3.5-VL 8B, video-only achieves the strongest classification for Chinese (Task 1 Macro F1 61.5 ±6.0), while video + transcript performs best for US (71.5 ±5.4). The advantage of video over transcript-only is significant for both cultures (p < 0.01 for US, p < 0.001 for CN); adding transcript to video significantly improves US performance (p < 0.001) but not Chinese. The authors speculate this may relate to lower transcription quality for Chinese and to low-context versus high-context communication differences.

  • Scaling model size does not reliably improve classification. Comparing 4B, 8B, and 14B Intern3.5-VL variants, the 8B model achieves the best F1 in both US and Chinese contexts, and the 14B model leads on evidence quality, though confidence intervals overlap. The authors suggest the bottleneck lies in culturally informed training data rather than raw model capacity.

  • Chinese norms in the benchmark skew toward fewer categories. The top norm categories in the benchmark are Expressing criticism (US 33.7%, CN 7.1%), Requesting information (24.2% / 17.8%), Admiration (8.8% / 8.6%), Greeting (8.8% / 5.6%), and Rejecting a request (3.3% / 3.3%).

Methodology in Plain English

Video collection. The team pulled 5–7 clips of 2–3 minutes each per show from YouTube and cut them into 15-second sub-clips, on the reasoning that a 15-second segment usually contains one distinct social norm. The US shows are Suits and The Office (workplace, drama and comedy) and Friends and The Big Bang Theory (informal). The Chinese shows mirror this structure: The Best Partner and Amazing Night (workplace) and iPartment and Home With Kids (informal). Comedy was chosen in part because humorous contexts yield more norm violations, for example through sarcasm.

Candidate generation. Gemini-2.0-Pro was prompted to pick a norm category, state an applicable cultural norm, describe the context and the norm subject, label adherence or violation, and give verbal and non-verbal evidence. The norm categories came from a Linguistic Data Consortium taxonomy subset — Thanks, Apology, Admiration, Greeting, Farewell, Requesting Information, Rejecting a Request, Granting a Request, Expressing criticism, Agreement, Disagreement — plus a Custom Category. This produced 511 unique candidate annotations for the US dataset and 500 for the Chinese dataset.

Human review. Three annotators per culture — screened for monocultural identity, country of residence, primary language, earliest language in life, and education, and trained on 20 instances first — validated or modified every field, with an Additional Explanation box for noting modifications or confirming agreement. Annotation agreement was measured with free-marginal multirater Kappa rather than traditional metrics because the adherence/violation distribution is skewed and would produce prevalence bias.

Benchmark filtering and tasks. Model evaluation used the consensus-filtered 542-instance benchmark. Task 1 is binary adherence/violation prediction given the video, a Whisper-large-v3 transcript, the norm category, and the specific norm. Task 2 adds generation of verbal evidence (quoted speech) and non-verbal evidence (gaze, gesture, distance, facial expression). Evidence quality was scored on correct predictions only, using GPT-5 as an LLM judge with a 5-point rubric that separates content perception (scores 1–2 missing cues, 3 identifying cues without reasoning) from cultural reasoning (scores 4–5 matching human annotations); where multiple reference explanations existed, the maximum judge score was taken.

Analysis. Results are reported as Macro F1 with variance for both tasks, and a hierarchical logistic regression predicts whether each model was correct on each item, with fixed effects for the model and controls for cultural context, adherence/violation, show, and norm category, plus random intercepts for item difficulty. A parallel linear model predicts verbal and non-verbal evidence scores. For the regression, the many norm categories were grouped into four clusters following Allan (1994, 1998): Statements, Invitationals, Authoritatives, and Expressives. Estimated marginal means were computed with the emmeans R package.

Why This Matters

This is, to the authors' knowledge, the first dataset and benchmark that explicitly evaluates cultural norm adherence versus violation in video. It shifts evaluation from object and action recognition toward pragmatic, culturally situated understanding, and it provides evidence that the failure mode is not simply model capacity.

Real-world applications:

  • Global content moderation and recommendation. Systems that judge whether behavior in uploaded video is acceptable need culture-specific norm reasoning rather than a single global standard.
  • Cross-cultural media localization and dubbing. Understanding whether a scene adheres to or violates local norms helps decide what needs adaptation for a different market.
  • Social robotics and embodied assistants. Robots operating in homes or workplaces must read gaze, gesture, posture, and distance the way locals do — exactly the non-verbal evidence channel where the paper finds models weakest.
  • Training data curation for multimodal models. The finding that scaling does not fix cultural deficits points toward investing in culturally diverse video data rather than larger models.

Industry relevance: Any company deploying video understanding across multiple markets faces the asymmetry this paper documents (57% higher odds of correct US adherence predictions than Chinese, OR = 1.57). The reported gap in candidate annotation quality for Chinese shows (42.35%–53.39% edits versus 18.28%–26.04% for US) is a direct warning that automated cultural annotation pipelines will underperform for under-represented cultures.

Future Directions

  • Broaden beyond two countries and beyond "country as a culture." The authors explicitly limit their claims to the US and China and acknowledge they had to treat each as monolithic despite large intra-cultural variation. Extending the framework to more regions and to subcultures within countries is the obvious next step.
  • Close the non-verbal evidence gap. Since models score significantly lower on non-verbal than verbal evidence, improving grounding of gaze, gesture, distance, and facial expression — or adding supervision targeted at those cues — is a concrete research target.
  • Move from scripted television to naturalistic video. The dataset relies on scripted TV as a cultural mirror. Collecting naturally occurring human interactions would test whether the findings hold outside scripted media.
  • Grow annotation scale and diversity. The authors note they could only recruit three annotators per country, and hope the work paves the way for larger-scale, more representative, longitudinal annotations using their framework.
  • Address pluralism and disagreement. The full dataset is being released so that disagreement and pluralistic annotation approaches can be studied, not just the consensus-filtered benchmark.

Target Audience

Researchers and engineers working on video-language models, multimodal evaluation, and culturally grounded AI; benchmark builders in computer vision and NLP who need a template for human-AI annotation protocols; and practitioners deploying video understanding systems across markets who need to know where and why cross-cultural performance degrades. Readers without a multimodal or statistics background will still follow the dataset construction and headline findings, but the hierarchical modeling sections require intermediate familiarity.

Authors’ abstract

As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce VideoNorms, a dataset of cultural norm annotations from popular US and Chinese TV shows annotated with adherence or violation labels and (non-)verbal evidence. Through a human-AI collaboration framework, each item was first annotated by a large VideoLLM, and then reviewed by at least three trained monocultural annotators with significant lived experience in the target culture, resulting in a dataset of over 3,000 human judgments. Human verification showed disparity in US and Chinese norm extraction performance, cautioning against fully automatic approaches cultures under-represented in training data. Hierarchical linear modeling analysis of $7$ open-weight VideoLLMs' performance revealed that: 1) models perform worse in Chinese compared to US, particularly for norm adherence prediction; 2) models have more difficulty in providing non-verbal evidence compared to verbal evidence for norm adherence/violation predictions. Ablation studies confirm video modality is indeed necessary for accurate performance, and scaling model size does not yield classification score improvements. Our findings and data contribute to culturally grounded video model training and evaluation.

Read the original paper