Research
UniH$^3$: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
Overview Research area: Medical image restoration (MedIR) with a single "all-in-one" universal deep learning model spanning multiple imaging modalities and degradation types. Technical level: Advanced

- arXiv
- 2609.11156
- Published
- 2026-09-10
- Authors
- Zhiwen Yang, Jiayin Li, Chengyu Liu, Hui Zhang, Bingzheng Wei, Yan Xu
AI summary
Overview
- Research area: Medical image restoration (MedIR) with a single "all-in-one" universal deep learning model spanning multiple imaging modalities and degradation types.
- Technical level: Advanced. The paper modifies transformer attention formulations, adds a momentum-updated memory bank, and introduces hierarchical uncertainty-based loss balancing; understanding it benefits from familiarity with transformers, cross-attention, and multi-task optimization.
- Scope: One sentence: UniH³ is a single framework for 2D and 3D medical image restoration that jointly models what medical images share (hierarchical homogeneity) and how they differ (hierarchical heterogeneity), validated on two newly assembled benchmarks, MedIR-2D-500K and MedIR-3D-3K.
What This Paper Is About
Existing all-in-one medical image restoration models concentrate on telling tasks apart — distinguishing modalities, degradation types, or scanners so each gets specialized processing — while largely ignoring the anatomical structure that medical images share within and across modalities. The authors argue this neglect makes training harder and generalization worse as task counts grow, and that prior work also handles only coarse inter-task differences, not finer intra-task variation caused by different scanners, centers, or patient demographics. UniH³ addresses both sides at once by distilling and retrieving homogeneous anatomical priors to guide restoration, and by balancing optimization conflicts at both the task and the sample level.
Key Contributions
- A unified hierarchical framework (UniH³) that simultaneously models inter-task and intra-task homogeneity and heterogeneity for all-in-one medical image restoration, covering both 2D and 3D settings.
- Hierarchical Homogeneity Memory (H²M), a memory bank with task-specific and task-shared slots that progressively distills intra- and inter-task anatomical priors from high-quality images during training and adaptively retrieves the most relevant priors for each input, plus Homogeneity-Guided Attention (HGA), an efficient mechanism that injects those priors into the restoration pipeline.
- Hierarchical Heterogeneity Balancer (H²B), a loss-balancing strategy that extends uncertainty-based weighting from a single per-task scalar to a task term plus a sample-specific correction, mitigating both inter-task and intra-task optimization conflicts.
- Two large-scale benchmarks: MedIR-2D-500K (509,200 2D LQ–HQ image pairs across seven 2D MedIR tasks) and MedIR-3D-3K (3,522 3D LQ–HQ volume pairs across three 3D MedIR tasks).
Main Findings
- 2D all-in-one performance: On MedIR-2D-500K, UniH³ averages 36.77 dB PSNR and 0.9045 SSIM across the seven tasks, surpassing the second-best method AdaIR (36.56 dB PSNR, 0.9026 SSIM) by 0.21 dB in PSNR. Per-task PSNR/SSIM for UniH³: PET 44.89/0.9883, CT 43.65/0.9368, MRI 39.55/0.9564, X-ray 36.88/0.9368, OCT 35.96/0.8921, Ultrasound 27.80/0.8179, Pathology 28.63/0.8035.
- 2D all-in-one efficiency: UniH³ uses 28.96 M parameters and 26.33 G FLOPs, compared with 35.59 M / 39.49 G for PromptIR and 28.76 M / 36.74 G for AdaIR.
- 3D all-in-one performance: UniH³-3D reaches an average of 45.29 dB PSNR and 0.9698 SSIM on MedIR-3D-3K (PET 49.77/0.9958, CT 45.37/0.9448, MRI 40.73/0.9689). The text states this exceeds the second-best method, Restore-RWKV-3D, by an average margin of 0.55 dB in PSNR; Table 4 lists Restore-RWKV-3D at 43.43 dB average PSNR and Spach Transformer at 43.81 dB average PSNR.
- 2D single-task performance: When separate models are trained per task, UniH³ averages 36.90 dB PSNR and 0.9058 SSIM, improving PSNR by 0.15 dB over the second-best method, MambaIR (36.75 dB PSNR, 0.9046 SSIM).
- 3D single-task performance: The text reports that UniH³-3D improves average PSNR by 1.48 dB over the second-best Spach Transformer across the three tasks; the contents of that table are not shown in the provided paper text.
- Both components matter: Disabling H²M or H²B each reduces performance relative to using both (36.66 dB PSNR / 0.9034 SSIM for H²M alone, 36.64 dB PSNR / 0.9033 SSIM for H²B alone, versus 36.77 dB PSNR / 0.9045 SSIM for both, against a baseline of 36.52 dB PSNR / 0.9023 SSIM).
- HGA beats alternatives: HGA achieves 36.77 dB PSNR / 0.9045 SSIM with 28.96 M parameters and 26.33 G FLOPs, versus SFT at 36.74 dB / 0.9041 with 37.46 M parameters and 36.99 G FLOPs, cross-attention at 36.67 dB / 0.9035 with 29.60 M / 27.43 G, and no HGA at 36.52 dB / 0.9023 with 27.17 M / 25.50 G.
- Both homogeneity levels help: Adding inter-task homogeneity alone gives 36.61 dB PSNR / 0.9031 SSIM, intra-task homogeneity alone gives 36.72 dB / 0.9038, and combining them gives 36.77 dB / 0.9045.
- Both heterogeneity levels help: Mitigating inter-task heterogeneity alone gives 36.70 dB PSNR / 0.9036 SSIM, intra-task heterogeneity alone gives 36.73 dB / 0.9041, and combining them gives 36.77 dB / 0.9045.
- Components transfer across backbones: Applying H²M and H²B to other transformer-based U-shaped backbones improves their averages: Uformer from 36.51 to 36.68 dB PSNR, Restormer from 36.49 to 36.67, PromptIR from 36.54 to 36.70, and AdaIR from 36.56 to 36.75.
- Memory retrieval behaves as intended: Visualized retrieval attention maps show tokens query mainly the task-shared slot and their own modality slot; two PET spine tokens share 8/10 of their top-score entries, PET and CT spine tokens share 2/10, while a PET spine token and a PET lesion token share 0/10.
- Uncertainty is finer-grained than prior balancing: The conventional uncertainty loss estimates one uncertainty per task, whereas the proposed loss gives each sample its own uncertainty within a per-task distribution, which the authors state captures both inter-task and intra-task variability.
Methodology in Plain English
The pipeline starts with a U-shaped encoder–decoder that takes a degraded low-quality image, projects it into shallow features with a 3×3 convolution, converts them to deep features through four levels of Homogeneity-Guided Transformer Blocks, and finally predicts a residual image that is added back to the input to produce the restored output.
Two ideas sit on top of that backbone. First, a memory system: during training, pairs of low-quality and high-quality images are turned into features, and a set of learnable "prototype" vectors uses cross-attention to pull clean anatomical structure out of the high-quality features. These distilled priors are written into a memory bank using an exponential moving average, with separate slots for each task and one shared slot for all tasks. At inference, the low-quality feature acts as a query that retrieves the most relevant stored priors, and these retrieved priors are fed into the network. This distillation step happens only during training and is discarded at test time.
Second, the attention mechanism that consumes the priors was redesigned so the reliable, clean prior carries more weight than the degraded observation. Instead of simply adding the prior to the value, the authors add identity-based terms that bias the attention map toward the prior and away from the low-quality value, then apply learnable channel-wise weights so the mechanism can fall back to ordinary self-attention. The implementation uses transposed self-attention following Restormer.
Third, for training stability, the loss is reweighted. Standard uncertainty-based balancing learns one uncertainty scalar per task; the authors model uncertainty as a task-level scalar plus a sample-level correction predicted by a lightweight block that looks at the low-quality input, the model's prediction (with a stop-gradient), and the ground truth.
Implementation specifics: for the 2D model, block counts per level are N₁=2, N₂=N₃=3, N₄=4, input channel dimension C=48, task count T=7, memory length L=128, EMA coefficient α=0.99, 128×128 patches, batch size 14, L1 reconstruction loss, Muon optimizer for 6×10⁵ iterations with learning rate annealed from 3×10⁻⁴ to 1×10⁻⁷ on a cosine schedule. The 3D variant, UniH³-3D, uses N₁=N₂=1 and N₃=N₄=5, C=16, an initial learning rate of 5×10⁻⁵, 64×64×64 patches, and batch size 6. Evaluation uses PSNR and SSIM.
Why This Matters
- Impact on research: The work reframes all-in-one medical image restoration around shared anatomy rather than task discrimination alone, and it contributes two named public benchmarks. The paper also reports that a single all-in-one model reaches performance comparable to state-of-the-art single-task MambaIR models across seven tasks, and that fine-tuning the trained all-in-one model per task yields further gains, suggesting value as a transferable pretrained backbone.
- Real-world applications:
- Multimodal clinical imaging workflows such as PET/CT and PET/MRI, where several restoration problems coexist and maintaining one model per task is inefficient to deploy and maintain.
- Denoising of low-dose or noisy acquisitions across CT, PET, X-ray, OCT, and ultrasound.
- Super-resolution of MRI and pathology images to improve visualization of fine structure.
- Deployment settings where a single model reduces maintenance burden compared with a fleet of task-specific models trained separately.
- Industry relevance: The paper emphasizes deployment and maintenance efficiency as a motivation for universal models, reports parameter and FLOP counts that put UniH³ in a competitive efficiency range, and releases code at https://github.com/Yaziwel/UniH3. The work comes from Beihang University, Tsinghua University, and an independent researcher.
Future Directions
- Covering more degradation types per modality: The authors state as a limitation that they focus only on the primary restoration task within each modality and do not cover other tasks or degradations that may occur in the same modality.
- Toward more universal models: The paper calls for pursuing broader MedIR models that would benefit clinical diagnosis and downstream tasks, which it cites as future work.
- Extending the stated benchmark scope: With MedIR-2D-500K and MedIR-3D-3K defined over seven 2D and three 3D tasks, expanding task and modality coverage within the same benchmark framing is an obvious extension the paper invites.
- Reconciling reported 3D margins with the comparison tables: The stated 0.55 dB margin over the "second-best Restore-RWKV-3D" does not match the ordering shown in Table 4 (Restore-RWKV-3D at 43.43 dB average PSNR versus Spach Transformer at 43.81 dB), and the 3D single-task table's contents are not shown in the provided text — clarifying these comparisons would help.
Target Audience
- Researchers in medical image restoration and all-in-one/multi-task image restoration who want a hierarchical alternative to task-discrimination-heavy designs.
- Practitioners building multimodal clinical imaging pipelines who need one deployable model rather than several task-specific ones.
- Machine learning engineers interested in memory-augmented attention, prior retrieval, or uncertainty-based multi-task loss balancing in domain-specific settings.
- Benchmark builders and dataset curators, since the paper contributes two named large-scale datasets and their composition.
- Beginners may still follow the high-level motivation, but the attention derivations and loss formulations make the full paper an advanced read.
Authors’ abstract
All-in-One medical image restoration (MedIR) aims to address diverse tasks across modalities and degradation types using a single universal model. Existing methods typically prioritize modeling inter-task heterogeneity (e.g., distinct data distributions and degradation types). However, they largely neglect the inherent homogeneity present in medical images, such as widely shared anatomical structures within and across modalities, which can be leveraged to ease model training and improve generalization. To this end, we propose UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity for all-in-one medical image restoration. Specifically, to comprehensively exploit homogeneity, we introduce a Hierarchical Homogeneity Memory (H2M) module that progressively distills intra- and inter-task homogeneity priors from high-quality images during training, and adaptively retrieves the most relevant priors tailored to the input for guided restoration. These retrieved priors are then injected into the restoration pipeline via an efficient Homogeneity-Guided Attention (HGA) mechanism. Furthermore, to comprehensively address heterogeneity, we design a Hierarchical Heterogeneity Balancer (H2B) that mitigates both inter- and intra-task conflicts during optimization, facilitating balanced and effective multi-task learning. Extensive experiments on two large-scale benchmarks, MedIR-2D-500K and MedIR-3D-3K, demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration. We hope this work establishes a strong benchmark and advances the development of general-purpose medical image restoration models. Code is available at https://github.com/Yaziwel/UniH3.