Research
Culturally-Aware Conversations: A Framework & Benchmark for LLMs
Culturally-Aware Conversations: A Framework & Benchmark for LLMs Overview Research area: Natural Language Processing / cross-cultural evaluation of large language models (cs.CL), with grounding in soc
- arXiv
- 2510.11563
- Published
- 2025-10-13
- Authors
- Shreya Havaldar, Sunny Rai, Young-Min Cho, Lyle Ungar
AI summary
Culturally-Aware Conversations: A Framework & Benchmark for LLMsOverview
Research area: Natural Language Processing / cross-cultural evaluation of large language models (cs.CL), with grounding in sociocultural theory and cultural psychology.
Technical level: Intermediate. The sociocultural framing and dataset construction are readable without deep technical background, but the evaluation protocol assumes familiarity with NLP benchmarking.
Scope: The paper introduces the first framework and benchmark designed to evaluate whether LLMs adapt their conversational linguistic style to cultural, situational, and relational context, and shows that five top LLMs largely fail to do so outside Western norms.
What This Paper Is About
Existing cultural benchmarks for LLMs mostly test factual knowledge about traditions and customs through trivia-style questions, which the authors argue does not reflect what actually happens when a user from one culture talks to a chatbot. The paper's goal is to build a framework and dataset that measures cultural adaptation at the level of conversational style — politeness, directness, insistence, pride, gratitude, self-disclosure — in realistic dialogue. It then uses that benchmark to evaluate today's leading models across eight national cultures.
Key Contributions
- A framework grounded in sociocultural theory that formalizes linguistic style as a function of three axes: Situation, Interpersonal Relationship, and Cultural Context. Developed with cultural experts (4 professors in cultural psychology, behavioral science, and communication at R1 universities, each with over a decade of culture research).
- A dataset of contextualized conversations built from this framework, containing 48 conversations and 240 possible responses, each annotated to represent 8 cultural perspectives (America, India, China, Japan, Korea, the Netherlands, Mexico, Nigeria).
- A new set of desiderata for cross-cultural conversational benchmarks: conversational framing, stylistic sensitivity, and subjective correctness (Table 1).
- An empirical evaluation of five current LLMs on the benchmark, showing they struggle with cultural adaptation in conversational settings and perform best on Western cultures.
Main Findings
- Models fail at cross-cultural adaptation overall. No evaluated model reaches high accuracy across all eight countries. The highest single cell is GPT-5-mini at 72.92% for the Netherlands; the lowest is GPT-5-mini at 43.75% for India.
- Western cultures are handled best. Country averages are highest for America (64.17%) and the Netherlands (63.75%), and lowest for India (49.17%). The authors flag this imbalance as concerning because LLMs deployed in non-Western contexts will be less likely to match local communication practices.
- Per-model performance varies by country. GPT-4.1 and Claude-4.5-Sonnet both score 70.83% for America; GPT-4.1 scores 47.92% for Korea, and Claude-4.5-Sonnet scores 45.83% for both India and Japan. Gemini-2.5-Flash reaches 64.58% for the Netherlands; Claude-3.5-Haiku scores 45.83% for Mexico.
- Full country-level averages: America 64.17%, India 49.17%, China 55.43%, Japan 52.08%, Korea 52.88%, Netherlands 63.75%, Mexico 55.83%, Nigeria 58.75%.
- Cultural norms differ by relational context, not just nationality. For example, the annotators indicated that in India it is more common to show gratitude in the workplace, while in a familial context communication is more expectation-driven, which the authors link to a strong sense of duty in Indian families.
- Annotator observations align partly with prior empirical work. The Netherlands favors directness and Japan is very polite; Nigerian culture is described as very insistent on the acceptance of food and gifts across all relational contexts; Americans tend toward more self-disclosure than any other culture, with the largest gap in professional and day-to-day relationships.
- Subjectivity is treated as a feature, not noise. Rather than forcing a single correct answer, the authors define an accepted range of style as μ ± 0.674σ, which corresponds to the 25th and 75th percentiles of a standard normal distribution and covers the central 50% of the underlying distribution. Correctness is judged by whether a model's selected response falls into that culture-specific range after rounding.
Methodology in Plain English
The researchers first consulted cultural communication experts to identify six conversational situations where the ideal response varies across cultures — such as offering and accepting food, and discussing personal accomplishments. For each situation, they identified the stylistic axis along which responses vary (for example, Insistence–Yielding for food, Pride–Shame for accomplishments). They also selected eight interpersonal relationships spanning familial, workplace, and day-to-day contexts.
Then they built the dataset in three stages. First, they prompted OpenAI's o3 to generate a concrete, contextualized scenario from a chosen situation and relationship pair. Second, they had o3 turn the scenario into a multi-turn conversation with a fixed opening turn and a set of five responses that all convey the same idea but differ along the relevant stylistic axis. All 240 responses were validated by the authors for stylistic range, and minor edits were made to approximately 30 responses to sound natural and realistic.
Third, they ran a user study in which 24 annotators from eight countries — recruited through volunteers at the authors' university and through Prolific — selected the most appropriate response given the norms of the culture they grew up in. They used 8 volunteers and 16 Prolific users, roughly 3 annotators per country, paying $20/hour with an average completion time of 42 minutes.
For evaluation, five models (Gemini-2.5-Flash, GPT-4.1, GPT-5-mini, Claude-3.5-Haiku, Claude-4.5-Sonnet) were given the country, situation, characters, first turn, and the five candidate responses, and asked to output only the number of the most appropriate response, using default temperature.
Why This Matters
Impact on research. The paper argues that factual cultural benchmarks do not generalize to the stylistic challenges of culturally sensitive communication, and it offers a concrete alternative: an evaluation paradigm centered on dialogue rather than trivia, on style rather than content, and on ranges of acceptable answers rather than a single gold label. It bridges cultural psychology and generative AI, providing a human-validated resource where many existing cultural benchmarks are large LLM-generated corpora validated only on a small subset.
Real-world applications:
- Customer service chatbots deployed globally, where politeness and directness expectations differ by user culture and by the workplace relationship involved.
- Tutoring systems, where the same feedback must be framed differently depending on whether it is directed at a student, a peer, or a subordinate.
- Therapy and personal-support agents, where self-disclosure and expressions of humility or pride carry different social meanings across cultures.
- Workplace and professional communication tools such as writing assistants, where gratitude versus expectation and directness versus indirectness determine whether a message reads as respectful or rude.
Industry relevance. The results show that current commercial models from OpenAI, Google, and Anthropic perform best on Western communication norms, which matters directly for companies deploying these systems in non-Western markets. The authors emphasize that "appropriate behavior" emerges from the interaction of culture, situation, and interpersonal relationship, so systems cannot be adapted by nationality alone.
Future Directions
- Multilingual extension. A large acknowledged limitation is that the dataset is entirely in English; the authors say future work should evaluate LLMs in all languages, noting they chose English for model QA skills, robustness, and reasoning capability, as well as for prompt engineering and manual validation by the authors.
- Larger-scale annotation. The study used only 3 annotators per country, limited by participant availability on Prolific, and the 8 chosen countries reflect areas with high concentrations of Prolific users. The authors call for larger studies that include underrepresented cultures.
- Richer notions of culture. The authors acknowledge they simplify culturally-aware communication to appropriate linguistic style, and note that communication in every culture is complex, dynamic, and multi-dimensional. They also use nationality and language as a proxy for culture, which they note does not perfectly align with culture.
- Mitigating dataset bias from LLM generation. The pipeline uses an LLM at two stages (scenario generation and conversation generation), and the authors caution that any inherent bias or fairness concern in the generating model may propagate into the dataset.
Target Audience
This paper is most useful for NLP researchers working on cultural evaluation, fairness, and dialogue systems; cultural and behavioral psychologists interested in operationalizing cultural dimensions computationally; and applied engineers and product teams building conversational agents for global audiences who need to know where current models fall short. It is also suitable for readers who want a concrete, human-validated alternative to trivia-style cultural benchmarks.
Authors’ abstract
Existing benchmarks that measure cultural adaptation in LLMs are misaligned with the actual challenges these models face when interacting with users from diverse cultural backgrounds. In this work, we introduce the first framework and benchmark designed to evaluate LLMs in realistic, multicultural conversational settings. Grounded in sociocultural theory, our framework formalizes how linguistic style - a key element of cultural communication - is shaped by situational, relational, and cultural context. We construct a benchmark dataset based on this framework, annotated by culturally diverse raters, and propose a new set of desiderata for cross-cultural evaluation in NLP: conversational framing, stylistic sensitivity, and subjective correctness. We evaluate today's top LLMs on our benchmark and show that these models struggle with cultural adaptation in a conversational setting.