Skip to content
AI.info

Research

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

Overview Research area: Medical computer vision and multimodal large language models (MLLMs), specifically anatomy-aware visual grounding for chest X-ray (CXR) interpretation. Technical level: Advance

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation
arXiv
2601.03191
Published
2026-01-06
Authors
Anees Ur Rehman Hashmi, Numan Saeed, Christoph Lippert

AI summary

Overview

  • Research area: Medical computer vision and multimodal large language models (MLLMs), specifically anatomy-aware visual grounding for chest X-ray (CXR) interpretation.
  • Technical level: Advanced. The paper assumes familiarity with vision-language architectures, contrastive learning, object detection decoders (DETR-style), LoRA fine-tuning, and radiology evaluation metrics.
  • Scope: The paper introduces AnatomiX, a two-stage anatomy-aware MLLM that first localizes thoracic anatomical structures and extracts their features, then uses an LLM to perform nine CXR tasks spanning grounding, report generation, visual question answering (VQA), and image understanding.

What This Paper Is About

Existing medical MLLMs handle chest X-ray tasks reasonably well overall, but they struggle with spatial reasoning and can fail to recognize which anatomical structure they are actually looking at. The authors argue this happens because current models perform grounding in a single implicit step, unlike radiologists who iteratively identify, localize, and evaluate each anatomical structure before concluding. AnatomiX addresses this by explicitly identifying anatomical objects and extracting their features before the language model performs any downstream task.

Key Contributions

  1. AnatomiX, an anatomy-aware grounded multimodal large language model for CXR interpretation, built around a two-stage process (anatomy perception, then language-model reasoning) rather than single-step grounding.
  2. State-of-the-art grounding performance: the paper reports SOTA results on diverse grounding tasks while maintaining on-par or better performance on report generation, VQA, and image understanding.
  3. Demonstrated robustness across different datasets and challenging settings, including horizontally flipped images (left ↔ right) and images with radiographic markers manually removed, plus ablation experiments validating each component.
  4. An Anatomy Perception Module (APM) that jointly learns global image features, bounding-box localization for N = 36 thoracic anatomical objects, object-level feature tokens, and contrastive alignment to textual descriptions, with a vector database used for retrieval at inference.

Main Findings

  • Grounding gains: In Table 1, AnatomiX achieves phrase grounding IoU 0.46 / mAP 0.35 and anatomy grounding IoU 0.73 / mAP 0.66, versus RadVLM's 0.39 / 0.30 (phrase) and 0.60 / 0.49 (anatomy), CheXagent's 0.33 / 0.24 (phrase) and 0.18 / 0.09 (anatomy), and MAIRA-2's 0.32 / 0.24 (phrase) and 0.35 / 0.24 (anatomy). The text describes this as up to 15% improvement in phrase grounding and over 25% in anatomy grounding.
  • Grounded diagnosis and grounded captioning (Table 1, reported as GD / GC): AnatomiX records BERTScore 0.63 / 0.65, ROUGE 0.60 / 0.56, METEOR 0.42 / 0.48, RadGraph-F1 0.58 / 0.50, and CheXbert-14-F1 0.54 / 0.78. The text describes up to 30% gains in grounded diagnosis and over 25% improvement in grounded captioning.
  • MAIRA-2 fails grounded diagnosis and captioning, with near-zero NLG and clinical scores (e.g., BERTScore 0.01 / 0.08, RadGraph-F1 0.00 / 0.02), attributed to it not being trained for spatial or region-specific input.
  • Robustness to flipped images: RadVLM drops sharply on laterality-sensitive anatomical objects, reaching an average overall IoU/mAP of 0.108/0.08, while AnatomiX shows no performance degradation and achieves 0.712/0.605. The paper attributes RadVLM's behavior to reliance on orientation cues.
  • Robustness without radiographic markers: After manual removal of markers such as text labels or AP/PA indicators on a subset of images, AnatomiX still localizes anatomical structures and phrases correctly. The authors note only a subset of samples was used because manual removal was chosen to avoid visual artifacts from automated editing.
  • Report generation (Table 2): AnatomiX scores ROUGE 0.53, BERTScore 0.38, METEOR 0.21, RadGraph 0.26, and CheXbert-14 F1 0.42, exceeding MAIRA-2 (0.43, 0.25, 0.12, 0.17, 0.45), Radialog (0.51, 0.35, 0.18, 0.24, 0.48), MedGemma (0.37, 0.29, 0.18, 0.20, 0.40), RadVLM (0.45, 0.27, 0.12, 0.19, 0.32), and CheXagent (0.32, 0.16, 0.06, 0.15, 0.31) on all metrics except CheXbert-14-F1, where Radialog (0.48) and MAIRA-2 (0.45) score higher. The paper notes both of those models contain approximately 1.5× more parameters than AnatomiX.
  • Image understanding (Table 3): AnatomiX reaches CheXbert-14 F1 0.85, IoU 0.31, mAP 0.20, tying CheXagent on classification F1 (0.85) and IoU (0.31) while CheXagent reports higher mAP (0.22). Other baselines score lower (e.g., RadVLM 0.43, 0.28, 0.12; MAIRA-2 0.00, 0.16, 0.01).
  • VQA (Table 4): AnatomiX scores 0.86 open-ended BERTScore / 0.86 CheXbert-14 F1 and 0.89 close-ended BERTScore / 0.95 CheXbert-14 F1, closely tracking CheXagent (0.86 / 0.87 and 0.90 / 0.97) and clearly above RadVLM (0.07 / 0.04 and 0.23 / 0.67) and MAIRA-2 (0.07 / 0.31 and 0.10 / 0.81). The authors attribute small gaps to keyword mismatches with ground truth, such as answering "yes" when the reference is "yes, pneumonia is present."
  • Metric sensitivity caveat: The paper notes that while NLG metrics reflect linguistic and clinical quality, the described keyword-mismatch effect can lower metric scores despite clinical correctness.
  • Ablation results are referenced in the paper's contribution list and discussion; the specific ablation numbers are in supplementary material not included in the provided content.

Methodology in Plain English

AnatomiX separates the problem into two stages.

First, an Anatomy Perception Module (APM) looks at the chest X-ray. An image encoder converts the image into patch embeddings. A DETR-inspired decoder combines those patches with N = 36 learnable object tokens, one per thoracic anatomical structure, so that each token focuses on one predefined object. A fully connected projector turns these tokens into predicted bounding boxes, trained with a combination of L1 and IoU losses using λ₁ = 5 and λ₂ = 2, the default values used in DETR. A feature extraction module then uses the object tokens as queries and the image patches as keys and values in cross-attention to pull out fine-grained per-object visual features, which are projected to a lower dimension.

Second, the module is trained so those visual features line up with text. Each anatomical region has an associated sentence describing its radiological findings, and a frozen sentence encoder (BiomedBERT) embeds these sentences into 768-dimensional vectors. Because findings in different anatomical regions often co-occur, the authors avoid a standard CLIP-style contrastive loss (which assumes one correct positive pair) and instead build a self-similarity matrix over the sentence embeddings, then optimize a KL-divergence-based soft contrastive loss between the projected visual features and this matrix. The APM loss is the sum of the bounding box loss and the contrastive loss.

At inference, the sentence encoder is replaced by a compact vector database holding unique sentences and their embeddings, organized as N independent sub-databases (one per anatomical object), built from the validation set of Chest-ImaGenome. Each anatomical object token retrieves its most similar sentence.

The second stage is the language model. Both the image embedding and the anatomical object tokens are projected into the LLM's embedding space. The LLM is based on the MedGemma-4b-it architecture (excluding its vision encoder), and its vocabulary is extended with N special tokens, one per anatomical object, plus four spatial grounding tokens. The APM outputs and the user prompt are combined into a structured multimodal prompt template, and the LLM generates the response.

Training proceeds in three steps on 4 NVIDIA H100 GPUs with 80GB memory using the AdamW optimizer. Step 1 trains the APM end-to-end for 30 epochs at a 1×e⁻⁴ learning rate with N = 36. Step 2 aligns the APM and LLM embedding spaces by unfreezing only the two projectors, trained for 2 epochs at 2×e⁻⁴ on the report generation dataset. Step 3 performs instruction tuning on all nine tasks for 3 epochs with LoRA, training the LLM plus the two projectors while the APM stays frozen. The APM was trained on over 237,000 samples from Chest ImaGenome, which extends MIMIC-CXR and provides localized information for 36 anatomical objects. Language model training used nine datasets derived from eight publicly available CXR datasets (MIMIC-CXR-JPG, VinDr-CXR, MS-CXR, PadChest-Grounding, SLAKE, MIMIC-CXR-VQA, RaDialog-Instruct, and Chest-ImaGenome), plus VinDr-Instruct and an Anatomy Grounding dataset constructed from VinDr-CXR and Chest ImaGenome.

Why This Matters

The paper argues that simply fine-tuning natural-image MLLMs on large medical datasets can create false spatial correspondences rather than genuine anatomical understanding, and that anatomy-oriented design is the key to accurate spatial reasoning in medical MLLMs. This matters most in settings where mistaking left for right, or misidentifying which structure a finding belongs to, has direct clinical consequences.

Impact on research: The work reframes medical grounding as an explicit, iterative anatomical modeling problem rather than a token-insertion trick, and provides a public code and pretrained model release at aneesurhashmi.github.io/anatomix, alongside measured comparisons against RadVLM, MAIRA-2, CheXagent, Radialog, and MedGemma.

Real-world applications:

  • Grounded radiology reporting that points clinicians to the specific region supporting each statement.
  • Interactive VQA assistants that answer region-specific questions about a CXR.
  • Verification of model outputs through visual overlays, so a clinician can inspect the evidence for a prediction.
  • Robust interpretation pipelines that do not depend on orientation markers or acquisition metadata being present or correctly placed.

Industry relevance: The paper emphasizes parameter efficiency, noting that two higher-scoring baselines on one metric contain approximately 1.5× more parameters than AnatomiX. Training on 4 NVIDIA H100 GPUs with 80GB memory, and use of LoRA, suggest a deployment path that is not limited to the largest compute budgets. The work was supported in part through Minerva computational and data resources at the Icahn School of Medicine at Mount Sinai, under CTSA grant UL1TR004419.

Future Directions

  • Extending anatomy-oriented architectures beyond chest X-rays to other modalities such as MRI, as the authors propose.
  • Reducing potential redundancy in the current multimodal prompt, which the authors identify as an improvement target for the architecture.
  • Moving from single-turn to multi-turn interactions, since this study focuses on single-turn setups and the authors state that multi-turn would enhance flexibility and applicability.
  • Broader validation of robustness claims: the marker-removal experiment used only a manually edited subset of samples because manual removal was chosen to avoid artifacts from automated editing, leaving open the question of how the model behaves on larger marker-free sets. The paper also notes that AOR introduced region-level information for CXR interpretation but is not yet publicly available for testing and comparison, so that comparison remains outstanding.

Target Audience

This paper is most useful to researchers and engineers building multimodal medical AI systems, particularly those working on grounded vision-language models, medical report generation, or radiology VQA. It will also interest clinical informatics teams evaluating whether a model's spatial claims can be trusted, and machine learning practitioners who want a concrete architecture-level alternative to token-based grounding. Readers need working knowledge of MLLM architectures, object detection losses, and contrastive learning to follow the methodology in detail, though the motivation and headline results are accessible to a broader medical AI audience.

Authors’ abstract

Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techniques improve overall performance, they often fail to establish a true anatomical correspondence, resulting in incorrect anatomical understanding in the medical domain. To address this gap, we introduce AnatomiX, a multitask multimodal large language model for anatomically grounded chest X-ray interpretation. Inspired by the radiological workflow, AnatomiX adopts a two stage approach: first, it identifies anatomical structures and extracts their features, and then leverages a large language model to perform diverse downstream tasks such as phrase grounding, report generation, visual question answering, and image understanding. Extensive experiments across multiple benchmarks demonstrate that AnatomiX achieves superior anatomical reasoning and delivers over 25% improvement in performance on anatomy grounding, phrase grounding, grounded diagnosis and grounded captioning tasks compared to existing approaches. Code and pretrained model are available at https://aneesurhashmi.github.io/anatomix

Read the original paper