Research
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
Overview Research area: Natural Language Processing — multilingual question answering, cross-lingual fact-checking, and contradiction detection. Technical level: Advanced. The paper assumes familiarit

- arXiv
- 2510.11928
- Published
- 2025-10-13
- Authors
- Lorena Calvo-Bartolomé, Valérie Aldana, Karla Cantarero, Alonso Madroñal de Mesa, Jerónimo Arenas-García, Jordan Boyd-Graber
AI summary
Overview
Research area: Natural Language Processing — multilingual question answering, cross-lingual fact-checking, and contradiction detection.
Technical level: Advanced. The paper assumes familiarity with polylingual topic modeling, dense retrieval (ANN/ENN, FAISS), LLM prompting, and NLI-based entailment labeling.
Scope: The paper introduces MIND (Multilingual Inconsistent Notion Detection), a user-in-the-loop pipeline that detects factual and cultural discrepancies between the knowledge bases underlying multilingual QA systems, evaluated on a bilingual English–Spanish maternal and infant health chatbot plus two out-of-domain datasets.
What This Paper Is About
Multilingual QA systems can give conflicting answers to the same question across languages, because the underlying knowledge bases differ in coverage, cultural framing, or institutional norms. The authors argue for fixing this upstream, at the data level, rather than hoping the model resolves conflicts at answer time. MIND aligns documents across languages in a shared topic space, generates questions from one language's corpus, retrieves evidence from the other, and asks an LLM whether the two answers entail each other, contradict each other, or diverge for cultural reasons.
Key Contributions
-
The MIND pipeline, a four-stage LLM-aided, user-in-the-loop system that aligns multilingual documents with polylingual topic modeling (PLTM) and then detects discrepancies across languages by comparing an answer's faithfulness in one language against its interpretation in another. The pipeline distinguishes four labels:
no_discrepancy(ND),contradiction(C),cultural_discrepancy(CD), andnot_enough_info(NEI). -
ROSIE-MIND, an English–Spanish QA benchmark labeled for factual and cultural discrepancies, built on the bilingual maternal and infant health knowledge base behind Rosie (Mane et al., 2023), and refined into ROSIE-MIND-v2.
-
The FEVER-DPLACE-Q controlled dataset (185 manually reviewed samples), built with gpt-4o from 50 REFUTES and 50 SUPPORTS FEVER-v1 claims, 50 D-PLACE ethnographic definitions remapped to cultural discrepancies, and 35 NEI samples, used for ablation studies of two-answer-per-question detection.
-
A generalization test on WIKI-EN-DE, derived from 600 aligned English–German Wikipedia page pairs, extending evaluation beyond English–Spanish and beyond the health domain, plus an open-source release of the datasets, models, and a package to run MIND on new datasets.
Main Findings
-
Retrieval: topic-based methods win, weighted ones win more. Under gpt-4o on topic t₁₆, TB-ENN-W (weighted, exact nearest neighbor) "achieves the highest retrieval performance overall," followed closely by TB-ENN. TB-ANN is slightly better than the ANN baseline across all metrics. ANN-based methods had the weakest ranking metrics but the lowest retrieval times (ANN: 0.015 ± 0.000 s; TB-ENN: 0.303 ± 0.007 s). Dynamically setting the threshold ε yielded only marginal time gains with no clear effect for gpt-4o.
-
Question generation is uniformly strong. Fulfillment was ≥ 94% across all six question criteria for the evaluated models. Annotator agreement was highest overall for llama3.3:70b (macro Gwet's AC1 0.88 on questions) versus qwen:32b (0.83) and gpt-4o (0.87).
-
Answer quality drops when moving from the anchor to the comparison corpus. Macro agreement for answers (anchor/comparison) was 0.85/0.74 for qwen:32b, 0.92/0.78 for llama3.3:70b, and 0.97/0.83 for gpt-4o. The authors attribute the drop to Spanish passages sometimes not fully aligning with the questions.
-
Discrepancy detection is strong but split by model strength. On FEVER-DPLACE-Q, weighted F1 was 0.918 for gpt-4o, 0.914 for qwen:32b, and 0.885 for llama3.3:70b. qwen:32b was best on
cultural_discrepancy(0.889) andno_discrepancy(0.980); gpt-4o led oncontradiction(0.962) andnot_enough_info(0.889). Inter-annotator agreement on the human evaluation was Fleiss's κ = 0.743. -
Category boundaries are unstable. In the human-vs-model comparison, gpt-4o failed to detect any discrepancy type, annotators disagreed on the
cultural_discrepancycases flagged by qwen:32b and llama3.3:70b, and cases MIND labeled as cultural discrepancies were often relabeled by annotators as contradictions or no-discrepancy. -
Regulatory and institutional variation drives most ambiguity. For "Are all newborns required to undergo cardiac screening tests?", the anchor answer says yes (standard newborn screenings) while the comparison says no, with California only requiring professionals to offer them; annotators called this a contradiction. A second example about impetigo (return after 48 hours vs. 24 hours) was labeled no discrepancy despite differing institutional policies.
-
Scale of detected discrepancies in ROSIE. Applying MIND (with llama3.3:70b as the final helper) to a 500-sample of anchor passages across topics t₁₂ (Pregnancy), t₁₆ (Infant Care), and t₂₅ (Pediatric Healthcare) detected 71, 136, and 50 contradictions; 35, 70, and 33 cultural discrepancies; 5,990, 4,313, and 3,937 no-discrepancy cases; and about 35K NEIs for t₁₂, t₁₆, and t₂₅ respectively.
-
Discrepancy type tracks topic. Contradictions were more frequent in domains with stronger medical guidelines (e.g., pregnancy), while cultural discrepancies prevailed in child development, likely reflecting differing parenting practices. The high number of NEIs suggests many English passages lack a direct Spanish counterpart.
-
Generalization holds outside health. On WIKI-EN-DE, English sources describe Anglo-American Freemasonry as requiring belief in a supreme being while German sources report that Liberal Freemasonry does not (a decision of the Grand Orient de France in 1877). Misaligned pages can also produce direct contradictions, such as the attribution of "The Preservation of St Paul after a Shipwreck at Malta" (English) versus "St Paul auf der Melite" (German).
-
False positives can come from the QA system itself, not the data. Some flagged discrepancies arise because generated answers omit details or blur nuance, meaning incomplete answers can make any LLM-based QA system look self-contradictory.
Methodology in Plain English
The authors treat the majority language corpus as the anchor and the other as the comparison corpus, without assuming documents are direct translations of each other. Both corpora are segmented into passages and fed to a polylingual topic model (MALLET's PLTM, K = 30, default parameters otherwise), which places documents from both languages into a shared thematic space even when they are only "loosely aligned." A human then discards noisy or off-domain topics.
For each remaining topic, the system takes anchor passages, generates yes/no questions grounded in that passage, and filters out subjective content and questions whose answers are not entailed by the passage (using an off-the-shelf NLI classifier). Each question is then decomposed into multiple reformulated search queries to handle contextual disambiguation. Those queries retrieve candidate passages from the comparison corpus using either exact or approximate nearest-neighbor search within topic clusters, with embeddings from BAAI/bge-m3 and FAISS IVF indices; retrieved passages can be scored with a topic-weighted similarity.
Answers are generated on both sides by prompting an LLM with the question and the corresponding evidence, with instructions to abstain when the passage lacks sufficient information. All questions are posed in the anchor language so that inconsistencies stem from the evidence rather than from translation drift. Finally, an LLM judges whether the comparison answer entails, contradicts, or culturally diverges from the anchor answer, and users review flagged cases. Evaluation used three LLMs (qwen:32b, llama3.3:70b, gpt-4o) with temperature 0 and top_p 0.1, plus crowdworkers on Prolific with graduate or doctorate degrees in Health & Welfare, paid £12/hour with at least three annotators per task.
Why This Matters
Impact on research. The paper reframes cross-lingual inconsistency as a data-level problem rather than a model-level one, and supplies a labeled bilingual benchmark (ROSIE-MIND) plus a controlled diagnostic set (FEVER-DPLACE-Q). It also shows that NLI-style three-way contradiction detection does not transfer cleanly to multilingual QA, where evidence sources may legitimately conflict, and it grounds the "cultural discrepancy" concept in Hymes' communicative competence, Spivak's notion of valid knowledge, and Fanon's epistemic asymmetries.
Real-world applications.
- Multilingual health chatbots such as Rosie, where conflicting guidance on breastfeeding, vaccination, or prenatal supplements has direct safety consequences.
- Cross-lingual encyclopedias and knowledge bases (the WIKI-EN-DE case), where the same entity or artwork is described differently across language editions.
- Content localization and QA knowledge-base curation teams who need to know where one language's documentation contradicts another's before deploying a retrieval system.
- Fact-checking and editorial tooling that surfaces candidate discrepancies for human reviewers rather than auto-correcting them.
Industry relevance. Any organization deploying retrieval-augmented QA in more than one language faces the failure mode this paper targets: authoritative sources in each language that disagree. The released repository provides the datasets, models, and a package to run MIND on new datasets, lowering the barrier to auditing a multilingual knowledge base.
Future Directions
- Finer-grained label categories. The authors suggest that adding a category such as "contextual differences" may capture cases that are not strictly cultural but do not fit cleanly as contradictions, given how often annotators relabeled MIND's cultural discrepancies.
- Incremental cross-lingual gap-filling. Because fully supervising large knowledge bases is infeasible, they propose an incremental strategy where documents in one language fill gaps in the other, organized by topic, with human oversight retained.
- Better context preservation at retrieval. The paper argues that QA systems like Rosie must preserve context during retrieval and avoid incomplete answers, since truncated answers create false discrepancies that look like data conflicts.
- Broader language and domain coverage. MIND applies to any language with loosely aligned or machine-translated corpora, provided multilingual embeddings exist; WIKI-EN-DE is presented as an initial extension, and use-case-specific annotators are recommended for scaling.
Target Audience
NLP researchers working on multilingual QA, retrieval-augmented generation, and cross-lingual fact-checking; practitioners building or auditing multilingual chatbots and knowledge bases, especially in health; and interdisciplinary teams in public health informatics or digital humanities who need to reconcile conflicting guidance across languages. The paper is most useful to readers comfortable with topic modeling and LLM-based pipelines, and it will be of particular interest to those who care about cultural nuance as a distinct category from factual contradiction.
Authors’ abstract
Multilingual question answering (QA) systems must ensure factual consistency across languages, especially for objective queries such as What is jaundice?, while also accounting for cultural variation in subjective responses. We propose MIND, a user-in-the-loop fact-checking pipeline to detect factual and cultural discrepancies in multilingual QA knowledge bases. MIND highlights divergent answers to culturally sensitive questions (e.g., Who assists in childbirth?) that vary by region and context. We evaluate MIND on a bilingual QA system in the maternal and infant health domain and release a dataset of bilingual questions annotated for factual and cultural inconsistencies. We further test MIND on datasets from other domains to assess generalization. In all cases, MIND reliably identifies inconsistencies, supporting the development of more culturally aware and factually consistent QA systems.