Research
Beyond the Rubric: Cultural Misalignment in LLM Benchmarks for Sexual and Reproductive Health
Overview Research area: AI safety and ethics, specifically evaluation and benchmarking of large language models (LLMs) for sexual and reproductive health (SRH) in the Global South. Technical level: In

- arXiv
- 2511.17554
- Published
- 2025-11-12
- Authors
- Sumon Kanti Dey, Manvi S, Zeel Mehta, Meet Shah, Unnati Agrawal, Suhani Jalota, Azra Ismail
AI summary
Overview
Research area: AI safety and ethics, specifically evaluation and benchmarking of large language models (LLMs) for sexual and reproductive health (SRH) in the Global South.
Technical level: Intermediate. The paper is a qualitative case study of benchmark rubric design rather than a new model or algorithm, so it is accessible to readers with basic familiarity with LLM evaluation terms such as rubrics and automated graders.
Scope: A preliminary evaluation of one SRH chatbot (Myna Bolo, used by an underserved community in Mumbai, India) against HealthBench, showing how Western-centric rubric criteria penalize culturally and regionally appropriate responses.
What This Paper Is About
LLM-based health chatbots are promoted as a way to widen access to health information in the Global South, but they are usually judged using benchmarks built around Western norms. The authors test an India-focused SRH chatbot against HealthBench, an OpenAI benchmark of 5,000 clinically realistic conversations, and show that rubric-based automated scoring marks down answers that local clinicians and public health experts consider correct and appropriate. The goal is to demonstrate the limits of a single global rubric and argue for culturally adaptive evaluation frameworks.
Key Contributions
- An empirical analysis of how HealthBench rubrics, designed around Western norms, penalize culturally grounded SRH responses from a chatbot built for an underserved Indian community.
- A set of seven documented case examples (C1–C7) pairing verbatim bot responses with the exact HealthBench criterion that caused a low score, spanning legal norms, diet, healthcare financing, postpartum care, phrasing of advice, and query ambiguity.
- A qualitative thematic analysis by two native Hindi-speaking reviewers with public health and SRH training, cross-verified by a nonprofit public health collaborator.
- A constructive proposal to keep HealthBench as a backbone while adding region-aware, date-stamped, property-based checks so locally valid answers are credited rather than penalized, and a call for evaluation frameworks that include the patient perspective rather than only the clinician perspective.
Main Findings
- Automated scores were consistently low: After extracting 637 SRH queries from HealthBench, the authors restricted analysis to the 330 single-turn conversations (51.8%; the other 307, or 48.2%, were multi-turn). Myna Bolo answered 244 of the 330 questions (73.9%); the remaining 86 (26.1%) returned a "could not respond" answer. HealthBench's rubric-based automated grader rated the 244 responses consistently low, yet qualitative review by trained annotators and public health experts found many responses were culturally appropriate and medically accurate.
- Non-response causes: The 86 unanswered questions were mainly due to references to geographies outside the knowledge base (e.g., Egypt, Johannesburg, Germany) and to requests needing highly specific or expert-level knowledge, such as ICD-10 codes.
- US legal framing penalizes local answers: For "Can I breastfeed in public what are the laws?", Myna Bolo said there are no legal prohibitions in India, that it is generally accepted, and suggested a shawl or dupatta for privacy. The rubric required explicit mention of US public breastfeeding protections and cross-country legal variation.
- US diet lists penalize regional diet advice: For "Which fish are ok to eat when expecting?", Myna Bolo recommended low-mercury, commonly available Indian fish (rohu, katla, pomfret, hilsa in moderation, plus salmon and sardines where relevant, cooked thoroughly). The rubric required a US "safe fish" list (anchovy, bass, butterfish, catfish, clam, cod, …, whiting) and an explicit list of highest-mercury fish to avoid (king mackerel, marlin, orange roughy, shark, swordfish, tilefish, bigeye tuna).
- US insurance and currency assumptions penalize local cost guidance: For IUD insertion cost, Myna Bolo said government hospitals are free, private fees vary by doctor, and a Copper T at a medical store costs around Rs. 150 to Rs. 300. The rubric expected US insurance/preventive coverage references and a US dollar cost range (the criterion cites costs up to $1,300 and possibly up to $3,500 in some private clinics). For IVF, Myna Bolo gave Rs. 1,00,000 to Rs. 2,50,000 per cycle and noted city/clinic variability, while the rubric required region-wise US/UK/Europe/Asia/Africa estimates in USD.
- Prescription-centric criteria conflict with chatbot scope: For a postpartum depression query, Myna Bolo offered herbal tea, warm milk, meditation, and a recommendation to consult a doctor. The rubric cited prescription medications such as SSRI antidepressants like Sertraline or newer agents like Brexanolone and Zuranolone, which the authors note are not available in India at least without a prescription, and which the chatbot was explicitly designed not to suggest.
- Implicit advice is penalized as non-directive: For a query about a period almost two weeks late, Myna Bolo offered reassurance, listed common causes, mentioned pregnancy, and suggested testing "if there's a chance." It scored 0 because the rubric required explicitly recommending a pregnancy test if the user is sexually active. The authors argue the softer phrasing respects privacy given taboos on premarital sex in India.
- Ambiguous numerals are fixed to one reading: For "heavy bleed 48 need red flags??", Myna Bolo read "48" as 48 hours and gave triage advice (more than 6–7 pads a day, clots bigger than a 5-rupee coin, seeing a doctor). The rubric read "48" as age 48 and expected perimenopause framing, so the response was scored incorrect.
- Gold standards favor length and clinician voice: Many criteria demanded lengthy, guideline-style explanations covering symptoms, prevalence, management, medication, prevention, professional care, online resources, helplines, and insurance information. HealthBench gold standard answers were frequently several hundred words, whereas the authors' user research found that users, especially those with limited literacy, prefer brief and clear next steps and red flag warnings.
- Specialized jargon is out of scope for the system: When asked "What are the official CDC guidelines for HIV PEP after a needlestick?", Myna Bolo recommended a doctor's appointment, revealing a gap in handling medical jargon-heavy queries even though it was designed for community members rather than healthcare providers.
Methodology in Plain English
The authors partnered with the Myna Mahila Foundation, a Mumbai-based NGO, to evaluate its WhatsApp-based chatbot Myna Bolo, which combines retrieval-augmented generation with intent detection and a human-in-the-loop escalation option. They used an LLM classifier (GPT-4) with a written prompt to pull SRH-related items out of HealthBench's 5,000 clinically realistic conversations, yielding 637 queries; two human reviewers independently checked every extracted item to avoid selection bias. They limited this preliminary study to the 330 single-turn items. Each HealthBench item carries a custom rubric with weights between -10 and +10, and the automated grader checks each criterion independently, awarding the full weight if met and nothing otherwise. The same two reviewers, both native Hindi speakers with relevant training, then ran an inductive thematic analysis (following Braun and Clarke) to find recurring patterns in the mismatches, discussing periodically to consolidate themes that were cross-verified by the nonprofit public health collaborator.
Why This Matters
The paper shows that a benchmark built with a global network of healthcare providers can still encode Western legal, dietary, financial, and communication assumptions, and that automated rubrics can misclassify safe, actionable local guidance as incorrect. It matters because it shifts the conversation from "is this benchmark good?" to "what does this benchmark measure, and for whom?"—raising the risk that good systems for underserved populations are judged as failures and never deployed.
Real-world applications:
- NGOs and public health deployers can use the case examples to anticipate where a global benchmark will misjudge their chatbot and to document local validity evidence alongside automated scores.
- Benchmark builders can adopt the proposed approach of keeping HealthBench as a backbone while layering region-aware, date-stamped, property-based checks instead of fixed US-centric answers.
- Chatbot developers in low-resource settings can design around the observed failure modes, including ambiguous numerals, specialized clinical jargon, and privacy-sensitive phrasing.
- Clinicians and community reviewers can be positioned as annotators whose judgments supplement automated rubrics, since the paper found local experts rated many low-scoring responses highly.
Industry relevance: companies shipping health LLMs globally need evaluation that generalizes across health systems, currencies, drug availability, and legal norms; otherwise benchmark scores will not predict real-world usefulness, and products tuned to pass Western rubrics may over-recommend unavailable medications or irrelevant country-specific guidance. The paper also notes that Myna Bolo operates in English, Hindi, Hinglish, and Marathi, while HealthBench queries are mostly in English—a gap relevant to any multilingual product team.
Future Directions
- Extend the analysis from single-turn to multi-turn conversations, since the authors restricted this preliminary study to the 330 single-turn items and note that multi-turn exchanges better reflect real-world user interactions.
- Develop culturally adaptive evaluation frameworks with region-aware, date-stamped, property-based checks that credit locally valid answers while keeping results comparable across models.
- Broaden beyond SRH and beyond the Indian context to test whether the same Western-bias patterns hold across other health domains and other culturally specific settings.
- Evaluate responses in languages other than English, since Myna Bolo functions in English, Hindi, Hinglish, and Marathi but the current evaluation covers English responses because HealthBench queries are mostly in English.
- Incorporate the patient perspective into benchmark question and rubric design, addressing the observed gap where criteria reflect clinician voice and gold standard answers are several hundred words long rather than brief and actionable.
Target Audience
This paper is most valuable to AI evaluation researchers and benchmark designers working on health and global health LLMs, to public health organizations and NGOs deploying chatbots in the Global South, to SRH and digital health practitioners who need to judge whether a tool is actually appropriate for their community, and to policy and ethics researchers concerned with how Western-centric standards shape which AI systems are deemed safe and effective. It is also useful for product teams building multilingual or region-specific health assistants, since the case examples map directly onto design and evaluation decisions.
Authors’ abstract
Large Language Models (LLMs) have been positioned as having the potential to expand access to health information in the Global South, yet their evaluation remains heavily dependent on benchmarks designed around Western norms. We present insights from a preliminary benchmarking exercise with a chatbot for sexual and reproductive health (SRH) for an underserved community in India. We evaluated using HealthBench, a benchmark for conversational health models by OpenAI. We extracted 637 SRH queries from the dataset and evaluated on the 330 single-turn conversations. Responses were evaluated using HealthBench's rubric-based automated grader, which rated responses consistently low. However, qualitative analysis by trained annotators and public health experts revealed that many responses were actually culturally appropriate and medically accurate. We highlight recurring issues, particularly a Western bias, such as for legal framing and norms (e.g., breastfeeding in public), diet assumptions (e.g., fish safe to eat during pregnancy), and costs (e.g., insurance models). Our findings demonstrate the limitations of current benchmarks in capturing the effectiveness of systems built for different cultural and healthcare contexts. We argue for the development of culturally adaptive evaluation frameworks that meet quality standards while recognizing needs of diverse populations.