Research
T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
Overview Research area: Medical vision-language models (MVLMs), model merging, and backpropagation-free test-time adaptation for zero-shot medical image classification. Technical level: Intermediate —

- arXiv
- 2510.27265
- Published
- 2025-10-31
- Authors
- Raza Imam, Hu Wang, Dwarikanath Mahapatra, Mohammad Yaqub
AI summary
Overview
- Research area: Medical vision-language models (MVLMs), model merging, and backpropagation-free test-time adaptation for zero-shot medical image classification.
- Technical level: Intermediate — the paper assumes familiarity with CLIP-style image-text encoders, softmax outputs, entropy, and KL/JS divergence, but explains its core mechanism clearly.
- One-sentence scope: The paper proposes T³ (pronounced /tee:cube/), a test-time merging framework that dynamically blends a pretrained generalist CLIP with a fine-tuned modality expert using Jensen-Shannon divergence between their output distributions, and evaluates it across four medical modalities under in-domain, base-to-novel, and corruption shifts.
What This Paper Is About
Medical vision-language models come in two useful but incomplete flavors: a fine-tuned "expert" that is highly accurate on the data distribution it was trained on but degrades under modality or dataset shift, and a large pretrained generalist that is robust but lacks fine-grained, site- or modality-specific nuance. Existing model-merging methods use a fixed or globally optimized interpolation weight and were largely designed for natural-image benchmarks, so they rarely deliver consistent gains across medical modalities. The paper's goal is an unsupervised, test-time rule that decides per input (or per batch) how much to trust the expert versus the generalist.
Key Contributions
- T³ — a non-iterative, backpropagation-free test-time task-adaptive interpolation framework. It computes Jensen-Shannon divergence between the pretrained and fine-tuned output distributions and maps it through a scaled sigmoid into a merging coefficient, learning optimal batch-wise merging weights without the cost of full backpropagation.
- A benchmark and cross-evaluation protocol for model merging in medical imaging. The protocol spans in-domain MedMNIST evaluation, base-to-novel transfer via MediMeta, and realistic corruptions via MedMNIST-C (noise and digital pixelation) across four modalities.
- An empirical analysis justifying dynamic test-time merging. This includes a decision-quadrant analysis showing that combined confidence alone fails to separate agreement from disagreement, whereas JS divergence isolates high-confidence disagreements, plus evidence that mutual information I(x) correlates positively with the entropy ratio R(x).
- State-of-the-art results in an unsupervised setting. T³ and its batch-wise variant T³_B outperform fine-tuned experts and both static and dynamic merging baselines across four medical modality tasks, while maintaining computational efficiency.
Main Findings
- JS divergence beats entropy ratio as a merging signal. The paper argues that DaWin's entropy ratio R(x) conflates confidence with agreement and cannot detect cases where both models are confident but assign high probability to different classes; the mutual information formulation I(x) = ½[KL(p_pt ∥ p̄) + KL(p_ft ∥ p̄)] is zero when the models agree and grows when they disagree.
- Accuracy gains across modalities (ViT-B/16). T³_B reaches mean Top-1 of 61.36 in Cell Microscopy versus 13.55 for Static Merging and 14.15 for DaWin; 68.94 in Breast Imaging versus the best baseline's 66.18; and 46.61 in Retinal OCT versus the Pretrained Expert's 46.60. T³_B attains near-expert accuracy on BloodMNIST (98.66 vs. 98.68).
- The paper describes Fundoscopy as "superior or competitive," not uniformly best. There, T³_B's mean is 55.78 and T³'s is 56.65, while Mixup Merging reaches 57.05; T³_B does post the highest in-domain Fundoscopy accuracy at 59.25 versus the Expert's 58.75.
- Large robustness improvements in mCE. T³_B yields 44.42 mCE in Cell Microscopy versus DaWin's 99.03 and Static Merging's 99.77, and 68.55 in Breast Imaging versus DaWin's 79.84 and Static Merging's 97.91.
- Aggregate robustness table. Averaged over the mean OOD of four medical datasets, Sample-wise T³ records 58.05 accuracy and 71.9 error, and Batch-wise T³_B records 58.17 accuracy and 71.7 error, compared with 55.01 / 86.7 for the Expert and 44.48 / 89.0 for DaWin.
- Cross-backbone consistency. T³ achieves mean accuracy of 58.05%, 56.62%, and 57.80% on ViT-B/16, ViT-L/14, and ResNet-50 respectively, exceeding experts and competing methods.
- Cross-modality consistency. The paper reports roughly a 2-3× performance gain over DaWin across all domains, with only minimal variance in relative improvement despite modality-specific differences in absolute performance.
- Efficiency. Without precomputation, the sample-wise T³ variant costs O(3N) and the batch-wise variant O(3B); with precomputation, T³_B is O(1B), matching vanilla pretrained and expert models. Precomputed T³_B processes in 41.3 seconds versus DaWin's 124.7 seconds, while un-precomputed runs are reported as ≥3800 seconds for T³ and ≥1260 seconds for T³_B.
- Coefficient variability. Figure 1 shows that the X-entropy ratio coefficient X(x) varies strongly by modality and shift type — for example, it stays tightly clustered for Fundoscopy in-domain but varies strongly under base-to-novel inputs, indicating reduced reliance on the fine-tuned expert.
Methodology in Plain English
The setup uses two fixed models of the same CLIP-style architecture: a pretrained generalist and a fine-tuned modality expert, both producing softmax distributions over the class labels for a given image. For each test input, the method measures how much the two distributions disagree using Jensen-Shannon divergence — essentially, an information-theoretic "do these two models see this case the same way?" score. That score is pushed through a sigmoid and scaled into an interpolation coefficient λ(x) between λ_min = 0.0 and λ_max = 1.0; a small value keeps the merged model near the generalist, a large value shifts it toward the expert. Two refinements are added: an extrapolation step (δ = 0.5 when either model's entropy falls below τ = 0.05) so that unusually confident predictions get a nudge, and a batch-wise variant that averages λ over all samples in a batch and performs a single parameter merge, cutting the number of interpolations from N to B.
The merged parameters are computed as θ_merged(x) = (1 − λ(x))θ_pt + λ(x)θ_ft — no labels, no gradient updates, no iterative optimization. Experiments are run in PyTorch on an NVIDIA A6000 48GB GPU using CLIP ViT-B/16 and ViT-L/14 backbones, with batch size BS = 32 and results averaged over three random seeds.
Why This Matters
Impact on research. The paper targets a gap it identifies explicitly: weight-interpolation methods were previously untested in medical VLMs under diagnostic test-time conditions, and prior medical model merging work focused on relatively smaller CNN architectures. By supplying both a divergence-based merging rule and a standardized cross-evaluation protocol spanning in-domain, base-to-novel, and corruption settings, it gives the field a shared benchmark for asking whether a merging method actually generalizes across medical modalities rather than helping on one.
Real-world applications.
- Cross-hospital deployment. A diagnostic model trained on one hospital's scanner, patient demographics, and imaging protocols must interpret scans from another institution; the paper's protocol directly simulates this via base-to-novel transfer on MediMeta.
- Degraded imaging conditions. Merging adapts when images arrive with noise or digital pixelation, the corruptions drawn from MedMNIST-C.
- Resource-constrained clinical environments. The framework is explicitly motivated by reducing memory and compute overhead relative to entropy-based test-time adaptation, which typically needs multiple augmentations and optimization.
- Modality-specific specialties. The four evaluated modalities — cell microscopy, breast imaging, fundoscopy, and retinal OCT — map onto distinct clinical reading workflows where specialist and generalist judgment would ordinarily be combined.
Industry relevance. Because the batch-wise variant with precomputed coefficients runs at the same O(1B) cost as a single pretrained or expert model (41.3 seconds versus 124.7 for DaWin in the reported comparison), the approach removes the usual accuracy-efficiency tradeoff that makes test-time adaptation hard to justify in clinical production pipelines. The paper states code is available at https://github.com/Razaimam45/TCube and that a full codebase and benchmarking setup will be released upon acceptance.
Future Directions
- Scaling beyond the evaluated backbones. Results are reported for ViT-B/16, ViT-L/14, and ResNet-50; whether the mutual-information criterion holds for larger or medically pretrained VLMs is not reported.
- Broadening the modality and task coverage. The benchmark spans four medical modalities; extension to other specialties, 3D volumes, or segmentation-style outputs is left open.
- Tuning the hyperparameters. The paper fixes λ_min = 0.0, λ_max = 1.0, δ = 0.5, and τ = 0.05 without reporting a systematic sweep, so the sensitivity of these choices across modalities remains an open question.
- Understanding the Fundoscopy exception. T³'s mean in Fundoscopy (56.65) is slightly below Mixup Merging (57.05), which suggests the conditions under which static merging still wins are not fully characterized.
Target Audience
Researchers and practitioners working on medical vision-language models, test-time adaptation, and model merging — particularly those interested in zero-shot or training-free deployment of CLIP-style models in clinical settings. It is also relevant to engineers building inference pipelines where compute budget and cross-site robustness both matter, and to benchmark designers looking for a cross-evaluation protocol that combines in-domain, base-to-novel, and corruption testing.
Note: the paper content provided is truncated mid-sentence in the Conclusion section, and several items the text references (Appendix B, Appendix D, Appendix E, Appendix 4, and Table 7) are cited but not included in the supplied content.
Authors’ abstract
In medical imaging, vision-language models face a critical duality: pretrained networks offer broad robustness but lack subtle, modality-specific characteristics, while fine-tuned expert models achieve high in-distribution accuracy yet falter under modality shift. Existing model-merging techniques, designed for natural-image benchmarks, are simple and efficient but fail to deliver consistent gains across diverse medical modalities; their static interpolation limits reliability in varied clinical tasks. To address this, we introduce Test-Time Task adaptive merging (T^3), a backpropagation-free framework that computes per-sample interpolation coefficients via the Jensen-Shannon divergence between the two models' output distributions. T^3 dynamically preserves local precision when models agree and defers to generalist robustness under drift. To overcome the inference costs of sample-wise merging, we further propose a batch-wise extension, T^3_B, that computes a merging coefficient across a batch of samples, dramatically reducing computational bottleneck. Recognizing the lack of a standardized medical-merging benchmark, we present a rigorous cross-evaluation protocol spanning in-domain, base-to-novel, and corruptions across four modalities. Empirically, T^3 sets new state-of-the-art in Top-1 accuracy and error reduction, outperforming strong baselines while maintaining efficiency, paving the way for adaptive MVLM deployment in clinical settings. Our code is available at https://github.com/Razaimam45/TCube.