Research
Enabling Real-Time Colonoscopic Polyp Segmentation on Commodity CPUs via Ultra-Lightweight Architecture
Overview Research area: Medical image segmentation, specifically real-time colonoscopic polyp segmentation for computer-assisted colorectal cancer screening, with a focus on CPU-only (GPU-free) deploy

- arXiv
- 2602.04381
- Published
- 2026-02-04
- Authors
- Weihao Gao, Zhuo Deng, Zheng Gong, Lan Ma
AI summary
Overview
Research area: Medical image segmentation, specifically real-time colonoscopic polyp segmentation for computer-assisted colorectal cancer screening, with a focus on CPU-only (GPU-free) deployment.
Technical level: Intermediate. The results can be read at face value, but the method section assumes familiarity with encoder–decoder segmentation networks, dilated/depth-wise convolutions, attention gating, and standard metrics (Dice, IoU, HD95).
Scope: The paper introduces the UltraSeg family of segmentation models (0.108M and 0.130M parameters, plus a 4.38M scaled variant), trained on mixed public polyp datasets and benchmarked against 12 comparison models for accuracy, cross-center generalization, and single-core CPU throughput.
What This Paper Is About
Polyp segmentation gives pixel-level outlines of lesions in colonoscopy video, but existing lightweight networks still carry around 10M parameters and need a GPU for real-time inference, which blocks deployment in primary hospitals, mobile devices, and embedded endoscopic systems. The authors start from a dermoscopy-derived ultra-light backbone (LB-UNet) instead of compressing a large network, and redesign it for colonoscopy so that the result runs faster than 30 FPS on a single commodity CPU core while keeping clinically usable accuracy. The goal is to establish the first strong baseline for the "extreme-compression" regime, defined in the paper as under 0.3M parameters with over 30 FPS on a commodity single-core CPU.
Key Contributions
- Establishes the extreme-compression paradigm (<0.3M parameters). UltraSeg-108K and UltraSeg-130K reach Dice scores of 0.7884 and 0.8038 under mixed-dataset training, with external validation on the unseen ETIS-Larib and BKAI-IGH datasets.
- Demonstrates intrinsic architectural efficiency across the capacity spectrum. Scaling from 0.13M to 4.38M parameters yields consistent gains, with UltraSeg-4.38M reaching 0.8503 Dice and surpassing the vanilla UNet-Base baseline using 14% of its parameters and 8% of its FLOPs.
- Validates real-time inference on commodity CPUs and memory-constrained edge devices, releasing not only model weights but also basic real-time video processing pipelines as a turnkey solution for resource-limited clinical settings.
- Introduces two new modules — the Enhanced Dilated Block (EDB) for multi-rate receptive field expansion and Cross-Layer Lightweight Fusion (CLLF) for cross-center, cross-modality generalization — plus a Predict-Gated Fusion module replacing hand-tuned fusion coefficients.
Main Findings
- Accuracy under extreme compression: At 256×256, UltraSeg-108K reaches 0.7884 Dice and UltraSeg-130K reaches 0.8038 Dice (HD95 37.31 pixels) on the mixed dataset, outperforming all other models below 0.3M parameters — LB-UNet (0.7685), EGE-UNet (0.7079), UNet-Tiny-DS (0.6716), and MobileUNet (0.4879).
- Parameter efficiency against a mid-size model: UltraSeg-130K delivers 96.6% of the Dice performance of UNet-Medium (7.76M parameters, 0.8318 Dice) while using 0.13M parameters.
- Zero-shot external validation: Without fine-tuning, UltraSeg-130K scores 0.6821 Dice on ETIS-Larib and 0.8001 on BKAI-IGH, versus 0.6797 and 0.7958 for UNet-Medium — i.e., it surpasses UNet-Medium on both external sets. Gaps to UNet-Base shrink to 2.0 and 1.6 percentage points.
- Comparison with a larger lightweight peer: UltraSeg-130K matches the accuracy of UNet-Light-DS (0.861M parameters, 0.8034 Dice) with substantially superior HD95 and stronger zero-shot generalization.
- Higher resolution helps: At 352×352, UltraSeg-130K reaches 0.8134 Dice on the mixed dataset, exceeding the 1.519M-parameter UNet-M-DS (0.8108), and reaches 0.6960 Dice on ETIS-Larib and 0.8013 on BKAI-IGH.
- CPU throughput: On a single-core Intel i5-14600K, UltraSeg-108K runs at 52.2 FPS at 256×256 and 30.8 FPS at 352×352; UltraSeg-130K runs at 51.7 FPS and 30.3 FPS respectively, versus 2.8 FPS and 2.3 FPS for UNet-Medium and 5.3 FPS and 4.7 FPS for UNet-M-DS.
- Cost of the design additions: The two cascaded EDBs at C=48 add 50k parameters while reducing FLOPs by 0.04 G relative to the same-stage vanilla bottleneck; CLLF adds under 0.022M parameters.
- Receptive field effect: The EDB yields a theoretical receptive field radius of 13 pixels on the feature map, equivalent to 52 pixels on the input image (encoder3 stride 4), covering roughly a 50-pixel diameter neighbourhood.
- Ablation of label generation: Replacing LB-UNet's genetic-algorithm boundary ground truth with a standard Canny edge detector (5×5 structural kernel) shortens boundary-label generation from hours to seconds with no degradation in segmentation accuracy.
- Reported gaps: Per-dataset intra-dataset results and the Kvasir-Instrument instrument-segmentation evaluation are referred to the supplementary document and are not reported in the main text. HD95 for UNet-Tiny-DS on BKAI-IGH is omitted due to excessive calculation error.
Methodology in Plain English
The authors reject the usual "top-down" recipe of taking a ResNet or ViT backbone pretrained on ImageNet and shrinking it, because that route still bottoms out around 10M parameters. Instead they work "bottom-up": they take LB-UNet, an ultra-light dermoscopy segmentation network under 0.1M parameters, and spend a small, deliberate parameter budget adapting it to colonoscopy. Their argument for borrowing from dermoscopy is that both dermoscopy and colonoscopy use visible-light surface imaging, both are single-region-of-interest segmentation tasks, and both have large public annotated benchmarks — the three conditions they say are required to develop native ultra-light models.
Concretely, they (1) reshape the channel schedule from LB-UNet's [8, 16, 24, 32, 48, 64] to [8, 16, 48, 64, 96], removing one downsampling stage and widening mid- and high-level channels to enlarge the receptive field and strengthen deep semantics; (2) insert Enhanced Dilated Blocks in the third encoder stage, which split channels into three groups and run depth-wise convolutions with dilation rates 1, 2 and 3 in parallel before fusing with a 1×1 convolution, so the layer sees fine, medium and global context at once; (3) add Cross-Layer Lightweight Fusion at the encoder–decoder bottleneck, comprising Attention Guided Fusion (two pixel-wise spatial masks α and β that sum to 1, blending Stage-3 and Stage-4 features) followed by Simple Spatial Attention (channel-average → Sigmoid gate → learnable residual scalar γ initialized to 0.3); and (4) replace fixed fusion hyperparameters with the Predict-Gated Fusion module, where learnable scalar weights α and β weight the region and boundary probability maps against the skip-connected encoder features. Training uses a dual-task loss: BceDiceLoss for the final region prediction, boundary deep supervision at three decoder levels with weights 0.1/0.2/0.3, and intermediate region supervision at four decoder levels with weights 0.1/0.2/0.3/0.4.
Experiments pool CVC-ClinicDB, Kvasir-SEG, PolypGen and PolypDB — each split 80:20 into train and test and then merged into a single mixed training set and mixed test set — with 10% of training data held out for validation. Default resolution is 256×256, batch size 4, Adam at learning rate 3×10⁻⁴, 100 epochs, early stopping after 10 epochs without validation Dice improvement; each experiment runs three times with fixed seeds on a single NVIDIA 5090 GPU. External validation is done directly on ETIS-Larib (196 images) and BKAI-IGH (1,000 images) without fine-tuning. Speed is measured by continuously inferring 1,000 images in single-core mode on an Intel i5-14600K.
Why This Matters
Impact on research. The paper reframes lightweight segmentation research as two distinct paradigms — accuracy-first (relaxed, above 1M parameters) and extreme-compression (capped at under 0.3M) — and argues that the second has been largely unexplored. It shows that architectural design, rather than post-hoc pruning or distillation, produces gains that persist from 0.13M to 4.38M parameters, and it supplies a reproducible blueprint that the authors suggest extends beyond endoscopy.
Real-world applications.
- Screening and diagnostic colonoscopy in primary hospitals that lack GPU servers.
- Embedded and capsule endoscopic systems where compute and power budgets are tight.
- Mobile or handheld clinical devices for real-time video interpretation.
- Turnkey clinical validation using the released real-time video processing pipelines alongside the model weights.
Industry relevance. The work targets the deployment gap that keeps state-of-the-art segmentation algorithms out of routine practice: the paper notes that even the lightest published polyp segmentation networks carry roughly 10M parameters and require GPU acceleration, while colonoscopy demands over 30 FPS. A 0.13M-parameter, single-core-CPU model with a 0.6 MB checkpoint changes the hardware economics of endoscopy suites, edge appliances, and any product where the GPU is the cost or power bottleneck. The paper also reports inter-scanner variability across manufacturers (Olympus, Pentax, Fujifilm) and the coexistence of WLE, NBI, BLI and LCI spectral modalities as the clinical realities the model must survive, and its CLLF module addresses exactly that.
Future Directions
- Extending the blueprint beyond endoscopy. The authors explicitly position the work as a reproducible template for real-time medical AI beyond endoscopy, but no other modality is evaluated here.
- Per-dataset and instrument-segmentation analysis. Intra-dataset results on CVC, Kvasir, PolypGen and PolypDB, plus the Kvasir-Instrument evaluation, are deferred to the supplementary document; whether the same conclusions hold in those limited-sample and non-polyp settings is not reported in the main text.
- Testing the scaling claim further. UltraSeg-4.38M (0.8503 Dice) demonstrates one scaled point; how far the design principles continue to pay off, and where they saturate relative to large Transformer- and Mamba-based architectures, remains open.
- Clinical validation in practice. The paper releases weights and basic real-time video pipelines as a turnkey solution, but no prospective clinical or reader study is reported, and the multi-center, multi-modal skew in PolypGen and PolypDB is described as a formidable remaining generalization challenge.
Target Audience
Researchers and engineers working on lightweight or deployable medical image segmentation; clinical engineering and product teams building colonoscopy or endoscopy hardware with CPU-only or edge compute constraints; and practitioners in resource-limited clinical settings who need real-time polyp segmentation without GPU infrastructure. Readers interested in the broader question of where the performance floor lies under a strict parameter budget will also find the extreme-compression framing useful.
Authors’ abstract
Real-time polyp segmentation is essential for early colorectal cancer detection, yet clinical deployment remains blocked by GPU dependency. We introduce the UltraSeg family, a set of CPU-native segmentation models operating below 0.3M parameters. UltraSeg-108K (0.108M) establishes the extreme-compression frontier, while UltraSeg-130K (0.130M) integrates cross-layer lightweight fusion for enhanced multi-center generalization. The architecture replaces parameter-heavy components with grouped multi-rate dilated convolutions and attention-gated cross-layer fusion, achieving real-time throughput on a single CPU core (exceeding 50 FPS at 256*256 and 30 FPS at 352*352) without sacrificing clinical-grade accuracy. Evaluated on seven public datasets, UltraSeg-130K attains Dice scores exceeding 0.8 at both resolutions, substantially outperforming all existing sub-0.3M competitors. Notably, it approaches or exceeds UNet-Medium (7.76M parameters) on zero-shot external validations while using only 1.7% of its parameters, establishing the first strong baseline for CPU-native real-time polyp segmentation. When scaled to 4.38M parameters, UltraSeg achieves accuracy competitive with heavyweight state-of-the-art models while maintaining an order-of-magnitude parameter advantage, demonstrating that the proposed design principles yield intrinsic representational gains across the entire efficiency spectrum. By delivering the first clinically deployable, CPU-native real-time solution, this work provides an immediately usable tool for resource-limited settings and a reproducible blueprint for real-time medical AI beyond endoscopy. Source code is publicly available.