Skip to content
AI.info

Research

RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture

Overview Research area: Computer vision and medical image representation learning — specifically self-supervised pretraining of a chest X-ray image encoder for downstream radiology tasks, with report

RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
arXiv
2601.15891
Published
2026-01-22
Authors
Anas Anwarul Haq Khan, Mariam Husain, Pratik Jalan, Kshitij Jadhav

AI summary

Overview

Research area: Computer vision and medical image representation learning — specifically self-supervised pretraining of a chest X-ray image encoder for downstream radiology tasks, with report generation as the main benchmark.

Technical level: Intermediate. Readers need some familiarity with Vision Transformers, masked-image modeling, contrastive vision-language pretraining (CLIP-style), and encoder-decoder report generation, but the paper's core idea (predict masked image regions in representation space rather than aligning images to text) is explainable without deep math.

Scope (one sentence): The paper adapts the I-JEPA latent-prediction framework to chest radiography, pretrains a ViT-B/14 encoder on 839,364 unlabeled X-rays, and empirically tests whether that language-free encoder transfers to report generation, classification, and segmentation.

What This Paper Is About

Most strong medical image encoders are trained by matching images to their paired radiology reports, which means they depend on text data and can absorb the biases of clinical writing — reports emphasize some findings and omit others. This paper asks whether an encoder trained with no language at all, by predicting the latent representations of masked regions from surrounding context, can still serve as the visual backbone for radiology report generation.

The goal is not a new architecture but a rigorous empirical test: freeze the RadJEPA encoder, bolt on a small trainable projector and a language decoder, and see how it compares against established image-text and image-only baselines across MIMIC-CXR and IU-Xray.

Key Contributions

  1. A chest X-ray adaptation of I-JEPA. The authors adapt the I-JEPA latent predictive pretraining protocol to chest radiographs and pretrain a ViT-B/14 encoder on 839,364 images drawn from five open-source datasets, with no report or label supervision.

  2. An extensive evaluation of language-free pretraining for report generation. A frozen RadJEPA encoder is coupled to a trainable two-layer projector and a Vicuna-7B decoder, and is also substituted into four established vision-language model pipelines, covering five language decoders in total.

  3. Controlled, data-matched comparisons. The paper includes control models (a RadJEPA control trained only on MIMIC-CXR, and a RAD-DINO control) to separate the effect of the predictive objective from the effect of domain-specific pretraining data, alongside comparisons against image-text, DINO-style, and JEPA-style baselines.

  4. Transfer beyond report generation. Complementary classification and segmentation experiments under common downstream protocols assess whether the encoder generalizes to other tasks. (Specific classification and segmentation numbers are not included in the provided paper content.)

Main Findings

  • RadJEPA leads on MIMIC-CXR report generation. With a frozen encoder and 256 tokens at 224×224 input, RadJEPA scores ROUGE-L 26.1 [25.7, 26.4], BLEU-4 10.1 [9.7, 10.4], RG ER 23.8 [23.4, 24.2], and Macro-F1-14 32.6 [31.2, 33.8]. The next best result on ROUGE-L among the listed baselines is RAD-DINO at 25.1 [24.7, 25.4], which uses 518×518 input and 1369 tokens.

  • RadJEPA also leads on IU-Xray. It scores ROUGE-L 28.4 [28.0, 28.7], BLEU-4 9.9 [9.6, 10.2], RG ER 27.5 [27.1, 27.9], and Macro-F1-14 27.6 [25.3, 29.8], ahead of RAD-DINO (26.3 ROUGE-L, 9.4 BLEU-4, 26.6 RG ER, 26.4 Macro-F1-14).

  • The predictive objective contributes beyond in-domain data. RadJEPA control (trained only on MIMIC-CXR, 197k images) reaches ROUGE-L 25.5 and Macro-F1-14 31.9 on MIMIC, and 27.1 / 26.8 on IU — above the RAD-DINO control (24.2 / 31.5 on MIMIC; 25.5 / 23.8 on IU), which was trained on the same 197k MIMIC-CXR images.

  • Smaller backbone, better results. RadJEPA uses a ViT-B/14 with 86M parameters, yet the authors report it consistently outperforms substantially larger image-only backbones — for example DINO-v2 (ViT-G/14, 1.1B parameters) scores ROUGE-L 22.7 on MIMIC and 25.4 on IU, and I-JEPA (ViT-H/14, 0.6B) scores 23.2 and 25.8.

  • Image-text baselines trail on clinical-label metrics. CLIP@224 reaches Macro-F1-14 of 24.7 on MIMIC (8.3 BLEU-4) and 18.1 on IU; CLIP@336 reaches 25.3 on MIMIC and 18.5 on IU. BioViL-T is the strongest image-text baseline on MIMIC Macro-F1-14 at 28.4, still below RadJEPA's 32.6.

  • CheXWorld is surpassed across all four report-generation metrics. CheXWorld scores ROUGE-L 24.1, BLEU-4 8.9, RG ER 20.1, Macro-F1-14 30.2 on MIMIC, and 26.1 / 9.1 / 26.3 / 24.9 on IU. The authors note CheXWorld's prior evaluation was limited and did not include report generation or AUPRC-based classification analysis.

  • Interpretation caveat stated by the authors. The broad comparisons reflect differences in pretraining data, model capacity, and input resolution as well as the pretraining objective; only the MIMIC-only controlled comparisons isolate the objective itself.

Methodology in Plain English

Pretraining. Take a chest X-ray. Cut out one "context" region and several separate "target" regions. A Vision Transformer encodes both. A predictor network looks only at the context representation and tries to guess the target representations. The loss is a SmoothL1 distance in representation space, with a stop-gradient on the targets, and the target encoder's weights are updated slowly from the context encoder via an exponential moving average. Nothing is reconstructed at the pixel level, and no text is used.

Data. Five open-source datasets are pooled: BRAX (41,620 images), CheXpert (224,316), MIMIC-CXR (300,491), ChestX-ray14 (112,120), and PadChest (160,817). MIMIC-CXR contributes only a subset of subjects to avoid overlap with evaluation sets. Public data skews roughly 6:1 frontal-to-lateral, so the authors add approximately 90K extra lateral MIMIC-CXR images to bring the ratio to roughly 3:1. Total: 839,364 images, 635,086 frontal and 204,278 lateral, from 372,854 subjects. Notably, RAD-DINO — the key image-only baseline — includes about 90K proprietary images, whereas RadJEPA uses only open-source data.

Downstream evaluation. The pretrained encoder is frozen everywhere.

  • Report generation: frozen visual embeddings pass through a two-layer MLP adapter built as a residual — the projected embedding equals the input plus a learnable scalar (initialized to 1) times a GeLU-activated two-layer transform. The projected tokens are concatenated with an instruction prompt and fed to Vicuna-7B, which generates the report autoregressively and is trained by negative log-likelihood. Only the adapter and decoder are updated.

  • Classification: a linear classifier on frozen features, trained with cross-entropy; only the head is trained.

  • Segmentation: multi-scale frozen features are aggregated by a UperNet decoder with a pixel-wise segmentation loss; only decoder parameters are optimized.

Why This Matters

Impact on research. The paper provides evidence that a purely visual, predictive pretraining objective can produce an encoder competitive at a language-heavy generative task, without consuming any report text during pretraining. It also isolates the objective from the data by including a data-matched MIMIC-only control, and it reports results at 224×224 with 256 tokens against baselines using up to 518×518 and 1369 tokens — a meaningful efficiency difference for anyone building clinical systems.

Real-world applications:

  • Automated radiology report drafting — a frozen encoder feeding a language decoder can generate preliminary impressions, letting radiologists edit rather than write from scratch.
  • Triage and prioritization — the same frozen features support disease classification heads (e.g., the 14-label Macro-F1 metric used here) for flagging studies.
  • Anatomical segmentation — the UperNet experiments indicate the features carry spatial information usable for outlining structures in chest X-rays.
  • Resource-constrained deployment — because language-free pretraining needs only images, it suits settings with large image archives but scarce or low-quality report pairs, and the 224×224 / 256-token configuration is lighter than 518×518 / 1369-token alternatives.

Industry relevance. Hospitals and imaging vendors hold far more unlabeled radiographs than clean image-text pairs, and text supervision is legally and operationally sensitive. A language-free encoder sidesteps dependence on report quality, avoids inheriting reporting bias from clinical narratives, and reduces token count per image — all of which lower the cost of building and serving clinical AI pipelines. The authors release code and pretrained weights.

Future Directions

  • Isolate objective versus scale more cleanly. The authors acknowledge that broad comparisons mix the pretraining objective with differences in pretraining data, model capacity, and input resolution. Larger-scale, matched-capacity comparisons would sharpen the conclusion.
  • Extend the encoder scale and resolution. RadJEPA uses a ViT-B/14 at 224×224; whether the advantage holds at higher resolution or with larger backbones is not answered here.
  • Broaden the clinical task set. Classification and segmentation results are reported only at a summary level in the provided content; fuller AUPRC-based classification analysis and additional anatomical targets remain open.
  • Push beyond chest X-rays. The paper notes that latent predictive architectures remain relatively underexplored in medical imaging, which invites adaptation to other modalities and body regions.

Target Audience

Researchers working on self-supervised and predictive representation learning who want a worked medical-imaging case study; medical computer vision engineers choosing a pretrained encoder for report generation, classification, or segmentation pipelines; and clinical AI teams evaluating whether they can build useful models from unlabeled image archives without paired reports. Readers interested in vision-language pretraining will find the controlled comparisons against CLIP, BioViL-T, BiomedCLIP, CheXzero, MRM, RAD-DINO, DINO-v2, DINO-v3, CheXWorld, and I-JEPA the most directly relevant section.

Authors’ abstract

Vision-language pretraining has driven progress in medical image representation learning, but it depends on paired image-text data and can inherit reporting bias from clinical narratives. We study whether language-free predictive pretraining can produce an image encoder that transfers effectively to radiology report generation. RadJEPA is a chest-X-ray adaptation of I-JEPA, pretrained on approximately 840K unlabeled radiographs using latent context-to-target prediction. Our primary contribution is an extensive empirical evaluation of this language-free encoder for report generation: the frozen image encoder is coupled to a trainable two-layer projector and language decoder, and is also substituted into four established vision-language backbones. Across MIMIC-CXR and IU-Xray, RadJEPA matches or exceeds the evaluated image-only and image-text baselines on lexical, entity-relation, and clinical-label metrics. Controlled MIMIC-only comparisons provide evidence that the predictive objective contributes beyond domain-specific pretraining, while broader comparisons also reflect differences in pretraining data, model capacity, and input resolution. Complementary classification and segmentation experiments assess transfer beyond report generation.

Read the original paper