Research
Language Models Entangle Language and Culture
Overview Research area: Multilingual evaluation of large language models, cross-lingual fairness, and the relationship between language and culture in LLM outputs. Technical level: Intermediate. The m

- arXiv
- 2601.15337
- Published
- 2026-01-20
- Authors
- Shourya Jain, Paras Chopra
AI summary
Overview
- Research area: Multilingual evaluation of large language models, cross-lingual fairness, and the relationship between language and culture in LLM outputs.
- Technical level: Intermediate. The methods (translation, LLM-as-a-Judge scoring, cluster analysis, non-parametric significance testing) are approachable, but the statistics and prompting setups assume some familiarity with LLM evaluation practice.
- Scope: The paper evaluates five LLMs across six languages on 20 advice-seeking questions and on a translated subset of CulturalBench, showing that answer quality and the cultural context of answers both shift with the language of the query.
What This Paper Is About
Users expect the same quality of advice from an LLM regardless of which language they type in. This paper tests whether that expectation holds, by asking the same set of real-world, advice-seeking questions (drawn from an analysis of the WildChat dataset) in English, Hindi, Chinese, Swahili, Hebrew, and Brazilian Portuguese, and then scoring the answers with an LLM judge. It goes beyond quality alone, investigating whether the language of a query also changes which cultural context the model draws on, and validating that with a translated subset of the CulturalBench factual knowledge benchmark.
Key Contributions
- A set of generic advice-seeking queries built from the authors' own analysis of the WildChat dataset, covering categories such as Healthcare, Business, Education, Research Advice, Trading/Investing, and Job/Interview preparation.
- What the authors describe as the first qualitative evaluation of LLM responses across languages for generic, advice-seeking queries, using culture-neutral prompts rather than prompts carrying cultural cues like name, nationality, or ethnicity.
- A translated version of the CulturalBench dataset (a subset of more than 750 questions from 29 regions, translated into Hindi, Chinese, Brazilian Portuguese, Swahili, and Hebrew), which the authors state they create and will release.
- Evidence that language and culture are entangled in LLMs, such that the language of the query changes the cultural context used in the response.
Main Findings
- All models perform best in English: Across every model evaluated, the highest answer quality appears in English, and each model produces worse responses in at least one language (Figure 2).
- Low-resource languages score lower: Responses in Hindi, Swahili, and Hebrew are consistently worse than those in English, Chinese, and Brazilian Portuguese. Even the Cohere-Aya models, which were trained for multilingual use cases, show worse performance in Swahili.
- Differences are statistically significant: Kruskal–Wallis tests give p-values below 0.05 for all models. Reported values are H = 712.7980 (p = 8.3941 × 10⁻¹⁵²) for aya-32b, H = 721.1299 (p = 1.3252 × 10⁻¹⁵³) for aya-8b, H = 610.8105 (p = 9.3325 × 10⁻¹³⁰) for magistral-small, H = 928.9057 (p = 1.4752 × 10⁻¹⁹⁸) for qwen3-14b, and H = 899.8367 (p = 2.8870 × 10⁻¹⁹²) for sarvam-m.
- The judge is not simply biased against non-English text: English responses translated into Hindi scored better than Hindi responses translated into English, indicating that the lower scores reflect genuinely lower-quality responses rather than the language in which the answer is written.
- Larger models are more consistent across languages: Cohere-Aya-32B shows smaller performance differences across languages than Cohere-Aya-8B.
- Fine-tuning shifts language performance: Although both Sarvam-m and Magistral are finetuned variants of Mistral-small-3.1-24B, they behave differently. Sarvam-m provides better responses for English and Hindi, while Magistral provides better responses for English, Chinese, and Brazilian Portuguese.
- Answers carry the culture of the query language: After translating all non-English responses into English, the LLM judge could still classify most answers as reflecting the culture associated with the language they were generated in (English/Western, Indian, Chinese, African, Latin American, or Jewish), as shown in Figure 3.
- CulturalBench accuracy varies by language: Qwen3-14b evaluated on the translated CulturalBench subset shows accuracy on questions about each country varying significantly across languages (Figure 4). The Kruskal–Wallis test gives H = 45.5158 and p = 1.1395 × 10⁻⁸.
- Language effects exceed random noise: Appending random strings to queries reduced performance but did not produce variation comparable to the language differences; the Kruskal–Wallis test across random strings gave H = 1.0228 and p = 7.9574 × 10⁻¹.
Methodology in Plain English
The authors started by cleaning the WildChat dataset: keeping only English queries, removing programming bug and error-fix queries (which dominate but come from a narrow set of users), keeping queries between 40 and 400 characters, and dropping duplicates or near-duplicates using the fuzzywuzzy library with a threshold of 60. They embedded the remaining queries with the Qwen3-0.6b embedding model, clustered them with HDBSCAN, and then manually analyzed the clusters to write 20 culture-independent advice-seeking questions. Those questions were translated into Chinese, Hindi, Brazilian Portuguese, Swahili, and Hebrew with Gemini-2.5-Flash at temperature 0.
Five models were tested: Qwen3-14B, Cohere-Aya-32B, Cohere-Aya-8B, Magistral, and Sarvam-m. Most were accessed through OpenRouter, with Cohere models and Sarvam-m called through their providers' own API platforms. For each question–language pair, 10 responses per model were generated at temperature 1.
Scoring used Cohere Command-A as an LLM judge at temperature 0. To validate the judge, the authors took 10 of their 20 queries and four languages (English, Hindi, Chinese, Hebrew), and had Cohere-Aya-32B generate five responses per query designed to correspond to scores 1 through 5 using a rubric covering detail and completeness, linguistic quality, factual correctness, actionability, and (in the judge prompt) riskiness. They compared six judge configurations; supplying 8 randomly chosen reference responses alongside the original query and response aligned best with ground-truth scores, as measured by Pearson correlation and Cohen's Kappa, so that configuration was used.
For the culture experiment, all non-English responses were translated into English (Gemini-2.5-Flash, temperature 0) and classified by the judge into one of six cultural categories. Separately, the translated CulturalBench subset was used to test Qwen3-14b at temperature 0 across all languages, including a control where random strings were appended to queries.
Why This Matters
The paper argues that a user's choice of language should not change the quality of advice they receive, since LLM responses influence decisions about healthcare, finance, and education. It provides direct evidence that this principle is violated in practice, and it adds a mechanism-level explanation: language and culture are entangled, so the language of a query silently changes the cultural assumptions baked into the answer.
Real-world applications:
- Consumer-facing LLM products: Teams serving multilingual users can use these findings to audit whether specific language groups get systematically weaker advice.
- Multilingual training data curation: Results point to pretraining data coverage per language and per culture as an explanation for the gaps, informing where to invest data collection.
- Benchmark development: The translated CulturalBench subset offers a reusable resource for measuring cultural knowledge across languages rather than only in English.
- Model selection for regional deployments: The Sarvam-m versus Magistral comparison shows that fine-tuning choices, not just base model family, determine which languages a model serves well.
Industry relevance: The findings matter for any organization deploying LLMs across regions, particularly those relying on "multilingual" models that are marketed as language-agnostic. The result that Cohere-Aya-32B is more consistent across languages than Cohere-Aya-8B suggests scale is one lever, while the divergence between two finetunes of the same base model suggests post-training is another.
Future Directions
- Extend the evaluation from small to moderate open-source models to larger models, which the authors list as a limitation.
- Apply mechanistic interpretability techniques to study how language and culture are represented inside LLMs.
- Improve multilingual training data and training methods to increase uniformity of response quality across languages, which the authors explicitly call for.
- Develop methods for identifying and mitigating language-based biases that disadvantage particular user groups.
- Address the reliance on LLM-as-a-Judge scoring, which the authors acknowledge may itself carry evaluation bias despite the absence-of-language-bias checks they ran.
Target Audience
This paper is most useful for researchers working on multilingual LLM evaluation, cross-lingual fairness, and cultural alignment; ML engineers and product teams building multilingual LLM applications; and benchmark designers interested in cultural knowledge and open-ended response quality. Readers primarily interested in benchmark accuracy scores on multiple-choice tasks will find the open-ended, culture-neutral framing here a deliberate departure from that tradition.
Authors’ abstract
Users should not be systemically disadvantaged by the language they use for interacting with LLMs; i.e. users across languages should get responses of similar quality irrespective of language used. In this work, we create a set of real-world open-ended questions based on our analysis of the WildChat dataset and use it to evaluate whether responses vary by language, specifically, whether answer quality depends on the language used to query the model. We also investigate how language and culture are entangled in LLMs such that choice of language changes the cultural information and context used in the response by using LLM-as-a-Judge to identify the cultural context present in responses. To further investigate this, we evaluate LLMs on a translated subset of the CulturalBench benchmark across multiple languages. Our evaluations reveal that LLMs consistently provide lower quality answers to open-ended questions in low resource languages. We find that language significantly impacts the cultural context used by the model. This difference in context impacts the quality of the downstream answer.