Research
Benchmark Success, Clinical Failure: When Reinforcement Learning Optimizes for Benchmarks, Not Patients
Overview Research area: Reinforcement learning for medical vision-language models, specifically R1-style training (supervised fine-tuning followed by Group Relative Policy Optimization, GRPO) applied

- arXiv
- 2512.23090
- Published
- 2025-12-28
- Authors
- Armin Berger, Manuela Bergau, Helen Schneider, Saad Ahmad, Tom Anglim Lagones, Gianluca Brugnara, Martha Foltyn-Dumitru, Kai Schlamp, Philipp Vollmuth, Rafet Sifa
AI summary
Overview
Research area: Reinforcement learning for medical vision-language models, specifically R1-style training (supervised fine-tuning followed by Group Relative Policy Optimization, GRPO) applied to multilabel chest X-ray classification.
Technical level: Intermediate. The paper assumes familiarity with supervised fine-tuning, LoRA, chain-of-thought reasoning, and policy-gradient reinforcement learning, but its central argument is accessible without deep technical background.
Scope: The paper trains a small vision-language model (ChexReason) on 2,000 SFT samples and 1,000 RL samples with a single NVIDIA A100 80GB GPU, and evaluates whether reinforcement learning improves diagnostic label prediction on CheXpert and NIH chest X-ray benchmarks, or only appears to.
What This Paper Is About
Reinforcement learning with verifiable rewards has markedly improved large language models on math and code, where correctness is easy to check. The authors ask whether the same recipe helps in a domain with far weaker supervision: multilabel chest X-ray diagnosis, where a model must produce both hard labels and reasoning traces. Their goal is to test whether R1-style training genuinely improves diagnostic reasoning, or whether it merely exploits the conventions of a single benchmark while losing the ability to transfer to other hospitals' data.
Key Contributions
-
Low-resource R1-style training: The authors present ChexReason, trained via SFT followed by GRPO with only 2,000 SFT samples and 1,000 RL samples on a single A100 GPU. They position this against NVIDIA's NV-Reason-CXR-3B, using 50 times less training data and 4 times less compute.
-
Instruction format sensitivity: Cross-model comparisons between MedGemma-4B (medically pre-trained) and Qwen2.5-VL-3B-Instruct (general-purpose) show that structured, medically informed reasoning scaffolds benefit general-purpose VLMs but give minimal gain for domain-specialized models.
-
Benchmark-transferability trade-off: GRPO improves CheXpert performance by 23% over the SFT checkpoint but degrades NIH transferability by 19%, mirroring the failures of the much larger NV-Reason-CXR-3B and suggesting a paradigm-level rather than scale-level problem.
-
Generalization paradox: The SFT checkpoint uniquely improves on out-of-distribution NIH data before any reinforcement learning is applied, indicating that teacher-guided reasoning traces capture more institution-agnostic features than reward-optimized outputs.
Main Findings
-
GRPO recovers in-distribution performance: On the CheXpert test set, ChexReason (SFT + GRPO) reached macro-F1 = 0.346, a 23% improvement over the SFT checkpoint (0.282) and close to the pretrained MedGemma baseline (0.362). Gains were concentrated in Cardiomegaly (0.442 to 0.664), Lung Opacity (0.161 to 0.743), and Support Devices (0.728 to 0.818).
-
Cross-dataset transfer collapses: On the NIH Chest X-ray test set, ChexReason fell to macro-F1 = 0.243, a 19% degradation from the SFT checkpoint (0.299) and a return to baseline MedGemma level (0.243). NV-Reason-CXR-3B showed the same pattern, dropping 61% from its CheXpert macro-F1 of 0.755 to 0.297 on NIH, while ChexReason dropped 30%.
-
The SFT checkpoint is the best NIH performer: Despite the weakest CheXpert test score among the compared variants (0.282), the SFT checkpoint scored 0.299 on NIH, improving over the MedGemma baseline, which itself dropped 33% on that dataset. The authors interpret this as teacher-guided traces capturing visual-semantic relationships that transcend specific label taxonomies.
-
The simpler reward function held up: The hard reward (format compliance plus Jaccard similarity) marginally beat the carefully engineered nuanced reward on CheXpert validation (macro-F1 = 0.258 vs. 0.257; micro-F1 = 0.391 vs. 0.387). Both improved macro-F1 over the SFT checkpoint (0.253 to 0.258) while micro-F1 declined slightly (0.415 to 0.391).
-
Prompt format effectiveness depends on medical pre-training: On base MedGemma-4B, the free-form Reasoning Narrative prompt scored best (micro-F1 = 0.524, macro-F1 = 0.270), above the structured Reasoning A baseline (micro-F1 = 0.498, macro-F1 = 0.245), and Reasoning C failed to decode in 48.2% of cases. After supervised fine-tuning, the ranking reversed: MedGemma-4B did best on direct label output (macro-F1 = 0.253 for Free Reasoning, 0.241 for Only Label), while Qwen2.5-VL-3B did best on structured 12-step Reasoning A (macro-F1 = 0.208, micro-F1 = 0.371).
-
Structured formats inflate token accuracy without reflecting learning quality: The authors observe that syntax-constrained variants (Only Label, Reasoning A, Reasoning Narrative) converge rapidly to high token accuracy, while Free Reasoning converges more slowly and saturates lower. They attribute this to output entropy and templated predictability rather than better learning.
-
GRPO appears to align to label conventions, not anatomy: The shared CheXpert label schema between MIMIC-CXR-JPG training data and the CheXpert test set, including identical pathology definitions, granularity, and uncertainty encoding, lets the reward signal exploit label-specific patterns that do not transfer to NIH's different labeling methodology.
Methodology in Plain English
The authors built a two-stage training pipeline. In stage one, they fine-tuned a vision-language model on 2,000 chest X-ray examples whose step-by-step reasoning traces were generated by Gemini 2.5 (given ground-truth labels but instructed to write as if reasoning independently). These traces were developed with radiologists from the University Clinic Bonn and Queensland Health, and a set of 100 datapoints was reviewed by radiologists to refine the prompting strategy. Training used LoRA adapters on the language model's attention and feed-forward layers with the vision encoder frozen, up to 6 epochs with early stopping, and loss masking restricted to assistant response tokens.
In stage two, they applied Group Relative Policy Optimization on 1,000 separate samples. GRPO samples a group of completions for each image and rewards those that score better relative to the rest of the group, avoiding the need for a separate value network. The authors used 4 completions per sample at temperature 0.8 with top-p = 0.95, a KL penalty of beta = 0.15 against the SFT checkpoint, Dr. GRPO loss normalization, and asymmetric clipping bounds of [0.15, 0.22]. They had to add stabilization mechanisms after observing early mode collapse. Two reward functions were tested: a "hard" reward combining strict output-format checks with Jaccard similarity between predicted and true labels, and a "nuanced" reward with multi-component scoring (100 points for exact match, partial credit for recall and precision, frequency-weighted penalties for false positives on common labels like "No Finding," and explicit penalties for repetition and mode collapse).
Separately, they ran a prompt-format ablation: nine prompt variants on base MedGemma-4B, then four of those formats as supervised fine-tuning targets on both MedGemma-4B and Qwen2.5-VL-3B-Instruct, using a 500-sample validation set drawn from the same MIMIC-CXR-JPG dataset.
Why This Matters
Impact on research: The paper argues that the failure mode is not a scale problem. NV-Reason-CXR-3B, trained with far more data and compute, degraded on NIH just as ChexReason did, so the authors conclude the issue likely lies in the RL fine-tuning paradigm when applied to small models on standardized benchmarks. That reframes a widely adopted recipe as one that may reward benchmark-specific shortcuts rather than diagnostic competence.
Real-world applications:
- Cross-institution deployment: Hospitals using different labeling conventions than the training institution could see diagnostic performance fall back to baseline despite strong benchmark scores.
- Low-resource clinical AI development: Teams without large annotation pipelines or GPU clusters can see that R1-style training is feasible on 2,000 SFT and 1,000 RL samples and a single A100, but should weigh the transferability cost.
- Regulatory and procurement evaluation: The results suggest that benchmark leaderboard position is an unreliable proxy for robustness across patient populations.
- Model selection for pre-training: The reversal between MedGemma-4B and Qwen2.5-VL-3B shows that whether to use structured reasoning prompts or direct label output depends on whether the base model already has medical pre-training.
Industry relevance: The paper directly implicates benchmark-driven development practices, arguing they may inadvertently disfavor models with broader real-world viability. It also gives a concrete data point that engineering a highly elaborate reward function did not outperform a simple one, which matters for teams allocating effort to reward design.
Future Directions
- Architectural or curriculum changes rather than post-hoc RL: The authors state that improving cross-dataset generalization may require architectural modifications, multi-dataset training curricula, or reward formulations that explicitly penalize overfitting to labeling conventions.
- Reward functions that resist single-institution overfitting: Since both hard and nuanced rewards aligned the model to CheXpert semantics without improving transfer, designing rewards that penalize convention-specific patterns is an open problem.
- Multi-institution training and evaluation: Training and testing across more than one labeling methodology would test whether the SFT checkpoint's NIH advantage (0.282 to 0.299) holds beyond a single pair of datasets.
- Reconciling benchmark gains with clinical robustness: Whether any R1-style configuration can deliver both the 23% CheXpert improvement and preserved cross-dataset transfer remains unresolved.
Target Audience
Medical AI researchers and engineers training vision-language models on chest X-ray or other radiology tasks; machine learning practitioners applying RLHF or GRPO-style methods in domains with subjective or weak supervision; clinical informatics and regulatory specialists evaluating whether benchmark performance predicts deployment robustness; and radiologists or clinicians collaborating on reasoning-trace datasets, who will recognize the institution-dependence problem the paper describes.
Authors’ abstract
Recent Reinforcement Learning (RL) advances for Large Language Models (LLMs) have improved reasoning tasks, yet their resource-constrained application to medical imaging remains underexplored. We introduce ChexReason, a vision-language model trained via R1-style methodology (SFT followed by GRPO) using only 2,000 SFT samples, 1,000 RL samples, and a single A100 GPU. Evaluations on CheXpert and NIH benchmarks reveal a fundamental tension: GRPO recovers in-distribution performance (23% improvement on CheXpert, macro-F1 = 0.346) but degrades cross-dataset transferability (19% drop on NIH). This mirrors high-resource models like NV-Reason-CXR-3B, suggesting the issue stems from the RL paradigm rather than scale. We identify a generalization paradox where the SFT checkpoint uniquely improves on NIH before optimization, indicating teacher-guided reasoning captures more institution-agnostic features. Furthermore, cross-model comparisons show structured reasoning scaffolds benefit general-purpose VLMs but offer minimal gain for medically pre-trained models. Consequently, curated supervised fine-tuning may outperform aggressive RL for clinical deployment requiring robustness across diverse populations.