Research
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
Overview Research area: Efficient vision backbones for supervised medical image classification (computer vision + medical AI). Technical level: Intermediate — the paper assumes familiarity with Vision

- arXiv
- 2510.27442
- Published
- 2025-10-31
- Authors
- Aon Safdar, Mohamed Saadeldin
AI summary
Overview
Research area: Efficient vision backbones for supervised medical image classification (computer vision + medical AI).
Technical level: Intermediate — the paper assumes familiarity with Vision Transformers, tokenization, self-attention, and benchmark evaluation, though the writing is accessible to readers who understand standard CNN/ViT terminology.
Scope: The paper introduces CoMViT, a roughly 4.5M-parameter Vision Transformer with a convolutional tokenizer, diagonal attention masking, learnable temperature scaling, and sequence pooling, evaluated on all twelve MedMNIST 2D datasets against larger CNN, AutoML, and ViT baselines.
What This Paper Is About
Vision Transformers perform well on vision tasks but are computationally heavy and tend to overfit when medical datasets are small, which limits their use in clinics with limited data and hardware. This paper asks whether a deliberately small, carefully redesigned ViT backbone — rather than a scaled-up or transfer-learned one — can match or beat much larger models across many medical imaging modalities. The authors build CoMViT and test it on the twelve MedMNIST 2D datasets to measure accuracy against parameter count and FLOPs.
Key Contributions
-
A compact ViT backbone for low-resource medical imaging. CoMViT combines a shallow convolutional tokenizer, a 7-layer transformer encoder, diagonal masking, dynamic temperature scaling, and learnable sequence pooling into a model of about 4.5M parameters.
-
Evidence that lightweight transformers can beat deeper models across modalities. The authors report that CoMViT matches or outperforms deeper CNN and ViT variants on the twelve MedMNIST 2D datasets while using 5 to 20 times fewer parameters.
-
Qualitative interpretability evidence. Grad-CAM analyses are used to show that CoMViT attends to clinically meaningful regions (for example cell boundaries in BloodMNIST, lesion areas in DermaMNIST, lung fields in PneumoniaMNIST, and relevant regions in TissueMNIST) despite its small size.
-
A public implementation. Code is released at the GitHub repository listed in the paper.
Main Findings
-
Average accuracy versus size: CoMViT reaches 84.5% average Top-1 accuracy on MedMNIST-2D with 4.55M parameters. For comparison, the paper reports ResNet-18 (11.2M) and ResNet-50 (23.5M) both at 0.821 average accuracy, auto-sklearn at 0.722, AutoKeras at 0.813, Google AutoML at 0.809, MedViT-T (10.2M) at 0.840, MedViT-S (23.0M) at 0.851, and MedViT-L (45.0M) at 0.842.
-
Per-dataset ranking: The paper states that CoMViT achieves the best or second-best accuracy on 8 of the 12 datasets. Its per-dataset accuracies in Table 2 span from 53.9 to 98.04 across the listed columns, which run Path, Chest, Derma, OCT, Pneumonia, Retina, Breast, Blood, Tissue, OrganA, OrganC, OrganS.
-
TissueMNIST head-to-head: Among "Tiny" models on TissueMNIST at 224×224, CoMViT reports 69.8% Top-1 with 4.55M parameters and 1.6 GFLOPs, compared with RVT-Ti (69.6%, 8.6M, 1.3G), ResNet-18 (68.1%, 11.7M, 1.8G), PVT-T (63.4%, 13.2M, 1.9G), PiT-Ti (62.1%, 4.9M, 0.7G), and DeiT-Ti (59.5%, 5.7M, 1.3G).
-
Competitive with larger "Small" and "Large" models: Still on TissueMNIST, Swin-T (29.0M, 4.5G) reaches 71.7%, Twins-SVT-S (24.0M, 2.9G) 72.1%, and MedViT-S (23.6M, 4.9G) 73.1%; among Large models, ResNet-152 (60.2M) is 67.5%, DeiT-B (87.0M) 66.9%, Swin-B (87.8M) 68.5%, and MedViT-L (45.8M) 69.9%.
-
Efficiency positioning: The authors describe CoMViT as sitting near the Pareto frontier of the model-size versus accuracy trade-off (Figure 4), where further increases in model size produce only marginal accuracy gains.
-
Parameter reduction claim: The abstract and introduction claim up to 5×–20× parameter reduction relative to the compared baselines without sacrificing accuracy.
-
Interpretability: Grad-CAM visualizations across six MedMNIST datasets are reported to highlight pathology-relevant regions rather than spurious artifacts, including on datasets with subtle or diffuse features such as TissueMNIST.
-
Ablation-style attribution: The discussion attributes CoMViT's gains to the combined effect of local bias from the convolutional tokenizer, attention refinement through diagonal masking and temperature scaling, adaptive pooling instead of a fixed class token, and a small encoder with a moderate embedding size — but no individual ablation study with separate numbers is reported.
Methodology in Plain English
Instead of cutting an image into fixed patches the way a standard ViT does, CoMViT first passes the image through a small convolutional "stem": two 7×7 convolution layers followed by a 3×3 max-pooling layer. This produces overlapping local feature tokens that already carry spatial context. Learnable positional embeddings are added to these tokens.
Those tokens then go through 7 transformer layers with 4 attention heads and a hidden size of 256, and an MLP block with 2× expansion (512). Two tweaks are applied inside attention: a diagonal mask that sets self-attention (a token attending to itself) to negative infinity, and a learnable temperature that scales the attention scores, which the authors say stabilizes gradients and encourages localized attention.
The output tokens are combined not with a [CLS] token but with learnable sequence pooling, where a softmax-weighted sum over tokens produces the final representation. The model is trained in PyTorch with timm using AdamW for 300 epochs, a learning rate starting at 1.1×10⁻⁴ with cosine decay, 10 epochs of warm-up and 10 of cooldown, RandAugment, Mixup (α = 0.8), CutMix (α = 1.0), label smoothing, drop-path 0.1, mixed precision, and gradient clipping at max norm 1.0. Batch size is 512 and inputs are 224×224; Mixup is probabilistically turned off after epoch 175. Evaluation covers Top-1 test accuracy, parameter count, and GFLOPs per forward pass, comparing against DeiT-Ti, PiT-Ti, PVT-Ti, RVT-Ti, ResNet-18, EfficientNet-B3, ResNet-50, AutoML systems, and MedViT variants.
Why This Matters
The paper argues that scale is not a prerequisite for strong medical image classification, and that tokenization plus attention design can substitute for raw model size. If that holds, it broadens who can build and deploy medical imaging models.
Research impact: It offers a lightweight, modality-agnostic baseline across all twelve MedMNIST 2D datasets, which may encourage further work on architecture design (rather than scaling or transfer learning) for small medical datasets, and it argues against relying on large natural-image pretrained ViTs that can suffer domain mismatch and negative knowledge transfer.
Real-world applications:
- Point-of-care or clinic-side screening tools where compute and memory are limited.
- Deployment on edge devices or modest hospital hardware for tasks such as chest X-ray or retinal image triage.
- Rapid prototyping and benchmarking of medical imaging pipelines on standardized data.
- Interpretability-conscious clinical decision support, where Grad-CAM maps can help clinicians see which regions drove a prediction.
Industry relevance: The 5×–20× parameter reduction and roughly 1.6 GFLOPs per forward pass on TissueMNIST make the model attractive for reducing inference cost, memory footprint, and energy use in medical AI products, and for on-premises deployment where sending patient data to large cloud models is undesirable.
Future Directions
- No dedicated ablation study is reported. The paper attributes gains to the combination of convolutional tokenization, diagonal masking, temperature scaling, and sequence pooling, but does not isolate how much each component contributes; measuring that is an obvious next step.
- Generalization beyond MedMNIST is untested. All twelve datasets are standardized to 224×224 low-resolution images under one benchmark; how CoMViT performs on full-resolution, multi-site, or 3D clinical data (such as CT or MRI volumes) is not reported.
- Non-classification tasks are unexplored. The work focuses on supervised classification; detection, segmentation, and ordinal or multi-label clinical tasks are not evaluated.
- Comparison scope could be widened. EfficientNet-B3 is named as a CNN baseline in the protocol, but no EfficientNet-B3 results appear in the reported tables, and no results are shown for the Tiny/Small/Large comparison outside TissueMNIST; extending these comparisons would strengthen the efficiency claims.
Target Audience
Researchers and engineers working on efficient or deployable deep learning for medical imaging; practitioners who need to classify small medical image datasets on limited hardware; and readers interested in Vision Transformer architecture design — particularly those focused on tokenization and attention mechanisms rather than scaling. It is also useful for graduate students using MedMNIST as a benchmark, since it provides a compact baseline with a released implementation.
Note: the paper's text is inconsistent in spelling the model name — it appears as both "CoMViT" and "ComViT" (and once as "CaMViT") — which is a presentation issue rather than a substantive one.
Authors’ abstract
Vision Transformers (ViTs) have demonstrated strong potential in medical imaging; however, their high computational demands and tendency to overfit on small datasets limit their applicability in real-world clinical scenarios. In this paper, we present CoMViT, a compact and generalizable Vision Transformer architecture optimized for resource-constrained medical image analysis. CoMViT integrates a convolutional tokenizer, diagonal masking, dynamic temperature scaling, and pooling-based sequence aggregation to improve performance and generalization. Through systematic architectural optimization, CoMViT achieves robust performance across twelve MedMNIST datasets while maintaining a lightweight design with only ~4.5M parameters. It matches or outperforms deeper CNN and ViT variants, offering up to 5-20x parameter reduction without sacrificing accuracy. Qualitative Grad-CAM analyses show that CoMViT consistently attends to clinically relevant regions despite its compact size. These results highlight the potential of principled ViT redesign for developing efficient and interpretable models in low-resource medical imaging settings.