Research
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark Overview Research area: Natural Language Processing — multilingual and domain-specific evaluation of Large Language Mod

- arXiv
- 2602.01714
- Published
- 2026-02-02
- Authors
- Mouath Abu-Daoud, Leen Kharouf, Omar El Hajj, Dana El Samad, Mariam Al-Omari, Jihad Mallat, Khaled Saleh, Nizar Habash, Farah E. Shamout
AI summary
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and BenchmarkOverview
Research area: Natural Language Processing — multilingual and domain-specific evaluation of Large Language Models (LLMs), specifically Arabic medical question answering.
Technical level: Intermediate. The paper is readable without deep ML background, but it assumes familiarity with benchmarks, zero-shot evaluation, few-shot prompting, and parameter-efficient fine-tuning (QLoRA).
Scope (one sentence): The paper introduces MedAraBench, a digitized Arabic medical multiple-choice question dataset with standardized training and test splits, validates it with expert clinician review and LLM-as-a-judge scoring, and reports zero-shot, few-shot, and QLoRA fine-tuning baselines for a set of proprietary and open-source LLMs.
What This Paper Is About
Arabic is among the most spoken languages in the world with over 400 million speakers, yet it remains underrepresented in medical NLP because open, expert-annotated resources are scarce. Most existing medical QA benchmarks (MedQA, MedMCQA, MMLU/USMLE) are English- or Chinese-centric, and the few Arabic resources that exist are limited in size, specialty coverage, difficulty mapping, or expert validation. The authors address this by building MedAraBench from scanned paper-based exams collected from regional medical school repositories, digitizing and cleaning them into a large MCQ benchmark, and using it to measure how well current state-of-the-art LLMs handle Arabic medical reasoning.
Key Contributions
- A large-scale Arabic medical benchmark. MedAraBench contains 24,883 MCQs spanning 19 medical specialties and five difficulty levels (Y1–Y5, corresponding to five years of study), with standardized training and test splits (19,894 training samples and 4,989 test samples).
- Two-layer quality assessment. The authors ran expert human evaluation with two board-certified clinicians (focusing on medical accuracy, clinical relevance, question difficulty, and question quality) plus an automated LLM-as-a-judge analysis over the full test set.
- Baseline benchmarking of 16 models. Proprietary models (claude-sonnet-4-20250514, gemini-2.0-flash, gpt-4.1, gpt-5, gpt-o3) and open-source models (deepseek-chat-v3-0324, qwen-plus, llama-3.3-70b-instruct, llama-3.1-8b-instruct, fanar-c-1-8.7b, allam-7b-instruct, c4ai-command-r7b-arabic-02-2025, medgemma-4b-it, apollo-7b, med42-8b, bimedix-bi) were evaluated zero-shot at temperature 0.
- Adaptation experiments and public release. Few-shot prompting and QLoRA fine-tuning were tested on llama-3.1-8b-instruct, and the dataset and evaluation scripts are released publicly.
Main Findings
- Proprietary reasoning models lead. gpt-o3 and gpt-5 achieved the highest overall accuracies at 0.765 and 0.764, followed by claude-sonnet-4-20250514 (0.694), gpt-4.1 (0.673), and gemini-2.0-flash (0.654).
- Open-source general-purpose models trail. deepseek-chat-v3-0324 reached 0.620 and qwen-plus 0.618; llama-3.3-70b-instruct reached 0.547, and llama-3.1-8b-instruct was lowest overall at 0.170.
- Arabic-centric and medical-specific models stayed below 0.5 accuracy. fanar-c-1-8.7b scored 0.498, allam-7b-instruct 0.447, c4ai-command-r7b-arabic-02-2025 0.381, medgemma-4b-it 0.390, bimedix-bi 0.390, med42-8b 0.318, and apollo-7b 0.238.
- Expert quality ratings were moderately high but agreement was only slight to fair. On the 378-question expert-reviewed sample: Medical Accuracy 0.722 [SD 0.448] with 82.0% agreement and Cohen's Kappa 0.555; Clinical Relevance 0.653 [0.476] with 65.6% agreement and Kappa 0.275; Question Difficulty 0.669 [0.471] with 65.6% agreement and Kappa 0.233; Question Quality 0.767 [0.423] with 68.3% agreement and Kappa 0.152.
- Sample size was determined statistically. Cochran's formula with z = 1.96, p = 0.5, and e = 0.05 yielded n₀ = 384 questions, adjusted to a final sample of 378 using the finite population correction.
- LLM judges only weakly track human experts. gpt-o3 had the highest Pearson correlations with Expert A and Expert B on Medical Accuracy (0.577 and 0.505), Clinical Relevance (0.252 and 0.377), and Question Quality (0.407 and 0.336). On Question Difficulty, alignment was weak to nonexistent, with gemini-2.0-flash highest for Expert A (0.019) and claude-4-sonnet highest for Expert B (0.039).
- LLM-as-a-judge metric averages varied by model. GPT-o3: Medical Accuracy 0.673, Clinical Relevance 0.827, Question Difficulty 0.588, Question Quality 0.841. Gemini 2.0 Flash: 0.717, 0.565, 0.815, 0.774. Claude-4-Sonnet: 0.711, 0.749, 0.576, 0.764. GPT-5: 0.533, 0.610, 0.597, 0.476.
- Fine-tuning beat prompting. For llama-3.1-8b-instruct, few-shot learning improved accuracy modestly by 12.4% (0.170 to 0.191), while QLoRA fine-tuning improved accuracy by 88.2% to 0.320, nearly doubling baseline performance.
- MedAraBench appears harder than MedArabiQ. Comparing generations, gemini-2.0-flash, gpt-4.1, gpt-5, gpt-o3, and qwen-plus performed better on MedArabiQ, while claude-sonnet-4-20250514 and llama-3.3-70b-instruct performed better on MedAraBench; all contemporary models outperformed legacy models on the MCQ task.
- The dataset is skewed toward early-year content. The Limitations section reports Y1 at 56%, Y2 at 22.89%, Y3 at 12.04%, Y4 at 3.38%, and Y5 at 5.14%. Appendix A reports different figures: Y1 15,095 questions (60.89%), Y2 4,954 (19.98%), Y3 2,313 (9.33%), Y4 1,033 (4.17%), Y5 1,396 (5.63%). The two sets of numbers in the paper do not match.
- Data filtering removed a large share of raw material. The initial pool was 34,333 MCQs; manual filtering reduced the dataset by approximately 29% to 24,883 samples. Nine questions with six answer choices were omitted.
- Specialty distribution is concentrated. Anatomy is the largest subset at 6,100 questions (24.61%) and Physiology second at 3,302 (13.32%); Pathology is smallest at 56 questions (0.23%). Other reported groups include Chemistry and Physics 3,894 (15.71%), Cell and Molecular Biology 3,155 (12.73%), Statistics 1,795 (7.24%), Biochemistry 1,387 (5.59%), Ophthalmology 1,318 (5.32%), Internal Medicine 762 (3.07%), Microbiology 750 (3.03%), Anesthesia 405 (1.63%), Surgery 357 (1.44%), Pharmacology 329 (1.33%), Embryology 184 (0.74%), and Emergency Medicine 120 (0.48%).
- Text length statistics. Appendix A.2 states the dataset has 24,791 questions with an average question length of 37.86 characters and average answer length of 171.09 characters; 18,554 questions have four answer choices and 6,228 have five. This question count differs from the 24,883 reported in the main text.
- Competitive positioning (Table 1). Compared against MedQA (60,000, English/Chinese), MedMCQA (193,000, English), MMLU-USMLE (1,800, English), MMLU Translation (15,000, 14 languages including Arabic), AraMed (270,000, Arabic), and MedArabiQ (700, Arabic), MedAraBench is listed at 24,000 MCQs and is the only row marked as Arabic, public, with expert annotation, difficulty mapping, and specialty coverage.
Methodology in Plain English
Data collection. The team gathered a large repository of scanned paper-based medical exams hosted on student-led social platforms of regional medical schools. Because these documents existed only as scans, professional typists manually digitized them. The material contained no personal or real patient data, so no anonymization was needed.
Cleaning. Five NLP researchers manually inspected and filtered the data, removing questions with missing or malformed correct answers, incomplete or duplicated choices, non-standard formatting, misaligned fields, ambiguous answer keys, or non-MCQ content. Each retained question carries annotations for number of answer choices (ABCD, ABCDE, or ABCDEF), difficulty level (Y1–Y5), and specialty inherited from the repository's original categorization. The authors note that the source material was not publicly available in structured digital form, which reduces the likelihood of data contamination.
Splitting. After dropping the nine six-choice questions, the data was split into a stratified random 80% training set and 20% test set so that each specialty is represented evenly in both. No external terminology standardization was applied, a choice the authors justify on the grounds that real-world medical QA is not necessarily standardized either.
Quality assessment. Two board-certified clinicians (Anesthesiology and Internal Medicine, each with over 20 years of experience and Arabic clinical fluency) independently rated a representative sample on four criteria — Medical Accuracy, Clinical Relevance, Question Difficulty, and Question Quality — on a high/low scale, using Qualtrics with double-blinded, pre-registered instructions. Question Quality decomposed into clarity, option homogeneity, single best answer, and no clueing. Separately, four top-performing LLMs were prompted to act as medical education experts and rate the entire test set on a binary (0 or 1) scale across the same four metrics; the paper's parenthetical list names only three of the four (gpt-03, gemini-2.0-flash, and claude-4-sonnet), while Table 3 reports four (GPT-o3, Gemini 2.0 Flash, Claude-4-Sonnet, GPT-5).
Benchmarking. Sixteen models were queried zero-shot on the test set at temperature 0, prompted to output only the answer letter, with responses parsed by pattern matching. No explicit language parameters were set in API calls because the models detect Arabic automatically. The paper's Section 4.4 refers to "the 15 evaluated models," which does not match the 16 rows in Table 4.
Adaptation experiments. For few-shot learning, three high-quality training-split questions rated highly by experts on all metrics — covering anatomy, biochemistry, and physiology — were supplied in Arabic as exemplars, kept out of the test set. For fine-tuning, QLoRA was applied to llama-3.1-8b-instruct loaded in 4-bit precision, with LoRA adapters on the key attention modules q_proj, k_proj, v_proj, and o_proj, training for up to 800 steps with batched gradient accumulation, and evaluated on the same test set.
Why This Matters
Research impact. The paper fills a documented gap in Arabic medical NLP resources and provides reproducible baselines on a benchmark that, by the authors' own comparison, appears more challenging than the earlier MedArabiQ benchmark. It also produces evidence that current LLM judges do not reliably agree with clinicians when assessing Arabic medical data quality, which is a caution for anyone planning to use automated evaluation pipelines.
Real-world applications:
- Clinical decision support in Arabic-speaking settings — the reported ceiling of 0.765 accuracy is well below expert-level performance, signaling that current models are not yet ready for clinical deployment.
- Arabic medical education and exam preparation — the Y1–Y5 difficulty mapping makes the dataset usable as a question bank tied to stages of medical training.
- Multilingual model development — the benchmark supplies a target for teams building or adapting models for underrepresented languages.
- Automated dataset curation — the LLM-as-a-judge correlation analysis informs how much automated quality screening can be trusted versus human review.
Industry relevance. The gap between proprietary and open-source model accuracy (top proprietary 0.765 versus many open-source and Arabic-centric models below 0.5) matters to organizations choosing between API-based and self-hosted models for Arabic-language health products, and the near-doubling of accuracy from QLoRA (0.170 to 0.320) matters to teams seeking to adapt smaller models to specialized domains rather than relying on prompting.
Future Directions
- Move beyond multiple-choice only. The dataset supports classification-style evaluation but not generative tasks; the authors suggest explanation-based tasks or clinician scoring of model justifications to test whether models reason or merely exploit statistical associations and lexical patterns.
- Expand modality and dialect coverage. The data is text-only, so specialties such as radiology and dermatology that depend on image-based reasoning are not covered, and the dataset assumes Modern Standard Arabic rather than the dialectal or mixed-language forms used in practice.
- Improve data quality and difficulty balance. The reviewers showed only slight to fair agreement on most metrics, and the data is skewed toward Y1 and Y2, so the authors call for higher-quality, more clinically relevant, and better-balanced benchmark datasets, and for broader consensus in future expert validation.
- Study better adaptation strategies. The authors propose exploring chain-of-thought prompting, additional few-shot configurations, and fine-tuning, as well as using Arabic lecture notes to extend medical Arabic NLP beyond MCQs.
Target Audience
This paper is most useful to NLP and clinical-AI researchers working on multilingual or low-resource language models, to medical AI teams needing an Arabic evaluation suite, to clinicians and medical educators involved in benchmark curation or in Arabic-language medical training, and to practitioners deciding whether current LLMs are trustworthy for Arabic-language clinical tasks.
Authors’ abstract
Arabic remains one of the most underrepresented languages in natural language processing research, particularly in medical applications, due to the limited availability of open-source data and benchmarks. The lack of resources hinders efforts to evaluate and advance the multilingual capabilities of Large Language Models (LLMs). In this paper, we introduce MedAraBench, a large-scale dataset consisting of Arabic multiple-choice question-answer pairs across various medical specialties. We constructed the dataset by manually digitizing a large repository of academic materials created by medical professionals in the Arabic-speaking region. We then conducted extensive preprocessing and split the dataset into training and test sets to support future research efforts in the area. To assess the quality of the data, we adopted two frameworks, namely expert human evaluation and LLM-as-a-judge. Our dataset is diverse and of high quality, spanning 19 specialties and five difficulty levels. For benchmarking purposes, we assessed the performance of eight state-of-the-art open-source and proprietary models, such as GPT-5, Gemini 2.0 Flash, and Claude 4-Sonnet. Our findings highlight the need for further domain-specific enhancements. We release the dataset and evaluation scripts to broaden the diversity of medical data benchmarks, expand the scope of evaluation suites for LLMs, and enhance the multilingual capabilities of models for deployment in clinical settings.