Research
A Women's Health Benchmark for Large Language Models
Overview Research area: Natural Language Processing / AI in healthcare — specifically the evaluation and safety benchmarking of large language models (LLMs) on women's health topics. Technical level:

- arXiv
- 2512.17028
- Published
- 2025-12-18
- Authors
- Victoria-Elisabeth Gruber, Razvan Marinescu, Diego Fajardo, Amin H. Nassar, Christopher Arkfeld, Alexandria Ludlow, Shama Patel, Mehrnoosh Samaei, Valerie Klug, Anna Huber, Marcel Gühner, Albert Botta i Orfila, Irene Lagoja, Kimya Tarr, Haleigh Larson, Mary Beth Howard
AI summary
Overview
- Research area: Natural Language Processing / AI in healthcare — specifically the evaluation and safety benchmarking of large language models (LLMs) on women's health topics.
- Technical level: Intermediate. The paper is a benchmark and evaluation study rather than a modeling paper; it requires no machine-learning mathematics, but familiarity with LLM evaluation concepts (prompts, approval/failure rates, confidence intervals) helps.
- Scope (one sentence): The paper introduces the Women's Health Benchmark (WHB), a 96-prompt expert-validated dataset spanning five medical specialties, three query types and eight error types, and uses it to evaluate 13 state-of-the-art LLMs, finding roughly 60% failure rates.
What This Paper Is About
Millions of people now ask AI chatbots health questions, and roughly one in six adults has used an AI chatbot for health-related questions in the last year, yet no benchmark existed that specifically measured how well LLMs handle women's health. Because women have historically been underrepresented in research and clinical trials, training data and model reasoning may contain sex- and gender-related gaps that produce confident but incorrect medical advice. The paper's goal is to build the first dedicated women's health benchmark (WHB) from expert-authored, clinically realistic prompts and to quantify how badly current LLMs fail on it.
Key Contributions
- The Women's Health Benchmark (WHB): 96 realistic open-ended prompts ("model stumps") across five medical specialties — obstetrics and gynecology, emergency medicine, primary care, oncology, and neurology.
- Expert-built dataset: Produced with 17 women's health experts, including clinicians, pharmacists and researchers across the United States of America and Europe. Experts prompted from a patient perspective, a clinician perspective, or asked evidence-based questions.
- A 13-model evaluation: The WHB was used to measure the performance of 13 LLMs, with breakdowns by query type, medical specialty and error type.
- Public release: The WHB data and code are released via The Lumos AI Labs Hugging Face repository (
TheLumos/WHB_subset, https://huggingface.co/datasets/TheLumos/WHB_subset).
Main Findings
- High overall failure: Current models show approximately 60% failure rates on the women's health benchmark, and no model demonstrated consistently high performance across all specialties, error types and query types.
- Best and worst models: The best performing model was GPT-5 with an approval rate of 53.1%, followed by Gemini 3 Pro at 47.9% and o3 at 46.9%. The worst performing model was Ministral-8B with an approval rate of 27.1%.
- Large versus small models: The mean approval rate for large LLMs (Claude Opus 4, Claude Sonnet 4, Gemini 2.5 Pro, Gemini 3 Pro, GPT-5, GPT-5.1, o3, Grok 4, Mistral Large) was 44%, versus 34% for small models (Gemini 2.5 Flash, GPT-4o Mini, Ministral 8B and o3 Mini).
- Missed urgency is a universal weakness: All models struggled with "missed urgency" indicators, which the authors say highlights the importance of human oversight for time-sensitive cases.
- Inappropriate recommendations improved in newer models: "Inappropriate recommendations" was the most variable error type, with GPT-5, GPT-5.1, o3 and Gemini 3 Pro showing the lowest failure rates, while GPT-4o-mini and Mistral-large-latest almost always failed on this error type. The authors caution that the sample size for this error type is relatively small (52 cases).
- Specialty differences: Neurology showed the highest failure rate at 76.9%, but this is based on the smallest sample size (39 cases total across all models). Oncology was second at 67.8%, emergency medicine third at 59.9%, and the lowest failure rates were primary care at 57.5% and obstetrics and gynecology at 56.7%. GPT-5 consistently performed better than other models across all specialties, followed by Gemini 3 Pro.
- Error type differences: The most common error type was "incorrect treatment advice" with a failure rate of 76.3%, followed by "outdated guidelines/treatment recommendations" at 69.2%. The least common error type was "missing critical information" with a failure rate of 49.5%.
- Query type differences: Models performed nearly equally on patient and clinician queries overall. GPT-5 showed the best performance on all query types, with failure rates of 47.1% (patient), 51.5% (clinician) and 33.3% (evidence/policy). Ministral-8B showed the worst performance on all query types, with failure rates of 78.4%, 57.6% and 91.7% respectively.
- Some models favor clinician queries: Grok 4, Ministral 8B, and Gemini 2.5 and 3 Pro showed dramatically better (that is, lower failure) performance on clinician queries compared to patient queries.
- Size and novelty matter: Overall, larger and newer models tended to perform better than smaller and older models.
Methodology in Plain English
A group of 17 experts, vetted for women's health expertise through literature screening and experience requirements, generated 345 realistic women's health questions over six weeks. Each question was randomly assigned to one of the 13 LLMs, and the expert judged whether the answer was correct. If the expert rejected the answer, they supplied a written justification (1–3 sentences) plus a supporting citation or link to an authoritative source, and the rejected prompt became a "model stump." Of the 345 questions, 249 (72.2%) produced answers classified as correct and 96 (27.8%) as incorrect. Experts contributed 143 model stumps with justifications and verified sources, of which 96 were approved after internal quality review; the rest were removed as duplicates, as not meeting women-specific criteria, or as lacking a factually incorrect response with sufficient clinical relevance or potential for harm.
The resulting 96 stumps were then used to benchmark all 13 models. A human evaluator with a PhD in clinical sciences reviewed each model answer and approved or rejected it, but could only reject an answer if it contained the exact same error identified by the expert group — an approved answer is not necessarily error-free, just free of that specific error. Failures were counted as incorrect cases over total cases, with 95% confidence intervals calculated using the Wilson score method.
The 13 models evaluated were Claude 4.0 Opus (2025-05-14), Claude 4.0 Sonnet (2025-05-14), Gemini 2.5 Flash (Preview 05-20), Gemini 2.5 Pro (Preview 05-06), Gemini 3 Pro, GPT-4o Mini (2024-07-18), GPT-5 (2025-08-07), GPT-5.1, Grok 4 (0709), Ministral-8B (Latest), Mistral Large (Latest), OpenAI o3 (2025-04-16), and OpenAI o3 Mini (2025-01-31).
Why This Matters
- Impact on research: The WHB is presented as the first evaluation framework for LLMs specifically in women's health, making a step toward addressing the under-representation of women's health in scientific research and health AI models. It argues that specialty-specific validation is essential.
- Real-world applications:
- Safety assessment and monitoring of consumer AI chatbots used by patients for women's health questions.
- Clinical decision support triage, especially for time-sensitive cases where "missed urgency" failures could cause harm.
- Regulatory and policy review of health AI, using expert-verified error categories as an evaluation standard.
- Dataset and model development, guiding construction of diverse, sex-aware training data.
- Industry relevance: Model developers can use the benchmark to compare versions and vendors; healthcare organizations can set expectations about human oversight and about which specialties and query types remain unreliable. The paper notes that women are especially likely to turn to LLMs for quick, accessible health information, and that inaccurate answers can worsen sex- and gender-related gaps in healthcare.
Future Directions
- Extending the benchmark to more medical specialties such as surgery, cardiology, and dermatology, and increasing the number of model stumps per specialty.
- Adding more query types, including diagnostic reasoning, treatment planning, and patient education.
- Replacing the single human evaluator with AI judges to reduce subjective interpretation of model outputs against predefined error categories and expert justifications.
- Developing a multi-turn benchmark to evaluate model performance on longer conversations.
Target Audience
Clinicians and women's health specialists, AI safety and evaluation researchers, LLM developers and model vendors, healthcare regulators and policy makers, and health informatics teams deploying patient-facing chatbots. The paper is also useful for readers seeking an introduction to medical benchmarking methodology, because it explains how expert-validated error taxonomies and approval/failure rates are constructed.
Note on reported figures: the paper does not report an overall numeric failure rate to more precision than "approximately 60%," and it notes that the neurology, oncology and "inappropriate recommendations" results rest on small sample sizes.
Authors’ abstract
As large language models (LLMs) become primary sources of health information for millions, their accuracy in women's health remains critically unexamined. We introduce the Women's Health Benchmark (WHB), the first benchmark evaluating LLM performance specifically in women's health. Our benchmark comprises 96 rigorously validated model stumps covering five medical specialties (obstetrics and gynecology, emergency medicine, primary care, oncology, and neurology), three query types (patient query, clinician query, and evidence/policy query), and eight error types (dosage/medication errors, missing critical information, outdated guidelines/treatment recommendations, incorrect treatment advice, incorrect factual information, missing/incorrect differential diagnosis, missed urgency, and inappropriate recommendations). We evaluated 13 state-of-the-art LLMs and revealed alarming gaps: current models show approximately 60\% failure rates on the women's health benchmark, with performance varying dramatically across specialties and error types. Notably, models universally struggle with "missed urgency" indicators, while newer models like GPT-5 show significant improvements in avoiding inappropriate recommendations. Our findings underscore that AI chatbots are not yet fully able of providing reliable advice in women's health.