Skip to content
AI.info

Research

Rectification Reimagined: A Unified Mamba Model for Image Correction and Rectangling with Prompts

Overview Research area: Computer vision, specifically low-level image processing — geometric distortion correction (portrait correction, wide-angle rectangling, stitched-image rectangling, rotation co

Rectification Reimagined: A Unified Mamba Model for Image Correction and Rectangling with Prompts
arXiv
2512.18718
Published
2025-12-21
Authors
Linwei Qiu, Gongzhe Li, Xiaozhe Zhang, Qilin Sun, Fengying Xie

AI summary

Overview

Research area: Computer vision, specifically low-level image processing — geometric distortion correction (portrait correction, wide-angle rectangling, stitched-image rectangling, rotation correction) within a single unified network.

Technical level: Advanced. The paper builds a mathematical general distortion model (Kannala-Brandt plus Brown-Conrady variants), uses Thin-Plate Spline warping, Mamba state-space blocks, and a Sparse Mixture-of-Experts layer, so it assumes familiarity with deep learning for image restoration.

Scope (one sentence): The paper proposes UniRect, a single prompt-conditioned Mamba-based framework that performs four different image correction and rectangling tasks either as four separate model instances or as one jointly trained model, and compares it against task-specific state-of-the-art methods on published benchmarks.

What This Paper Is About

Image correction and rectangling are handled today by separate, task-specific networks, which wastes memory and computation on devices such as smartphones that need several of these tasks at once. The authors argue that all four tasks are really inverse versions of one general lens-style distortion process, and that a single architecture can therefore learn them. Their goal is a task-agnostic model (called four-by-one) that can also be trained jointly across all four tasks (four-in-one) without the tasks interfering with each other.

Key Contributions

  1. A consistent rectification perspective and a general distortion model. The paper mathematically unifies the backward distortion processes of four photography tasks by combining the Kannala-Brandt model (used for portrait/wide-angle lens distortion) and the Brown-Conrady model (used for radial/stitched distortion) into one equation with distortion parameters, a rotation matrix, and a translation term.

  2. The Unified Rectification Framework (UniRect) with prompts. A task-agnostic architecture consisting of a Deformation Module and a Restoration Module, driven by visual prompts that indicate which task to perform, achieving the "four-by-one" setting.

  3. The Residual Progressive Thin-Plate Spline (RP-TPS) model. Instead of applying TPS once, two control-point predictors progressively predict deviations of control points, with all sampling performed from the original input to avoid intermediate interpolation degradation.

  4. A Sparse Mixture-of-Experts (SMoEs) structure for multi-task learning. A gating network with a top-k operator assigns weights across five expert networks to reduce task competition, enabling the "four-in-one" setting.

Main Findings

  • Portrait correction (T1): UniRect reports LineACC 66.523 and ShapeACC 97.454, compared with 66.143 / 97.253 for Shih et al. (2019), 66.784 / 97.490 for Tan et al. (2021), and 66.825 / 97.491 for Zhu et al. (2022). On the LLM-based metrics introduced because this dataset provides no image ground-truth, UniRect reports 7.168 LineACC-LLM and 8.152 ShapeACC-LLM versus 7.082 and 8.082 for Zhu et al. (2022).

  • Rectified wide-angle image rectangling (T2): UniRect reports PSNR 19.90 and SSIM 0.5721 against 18.68 and 0.5450 for RecRecNet and 13.90 and 0.3516 for ROP. However, its perceptual scores are worse than RecRecNet: FID 27.02 versus 19.01, and LPIPS 0.1245 versus 0.1136.

  • Stitched image rectangling (T3): UniRect reports PSNR 25.10, SSIM 0.7526, FID 19.59, and LPIPS 0.1120, compared with 21.28 / 0.7141 / 21.77 / 0.1557 for Nie et al. (2022), 14.70 / 0.3775 / 38.19 / 0.2846 for He et al. (2013b), and PSNR 20.42 / SSIM 0.6307 for MOWA.

  • Rotation correction (T4): UniRect reports PSNR 23.16, SSIM 0.7179, FID 6.55, and LPIPS 0.0873, compared with 22.29 / 0.6790 / 7.90 / 0.1970 for CoupledTPS, 21.02 / 0.6280 / 7.12 / 0.2050 for DRC, and 21.69 / 0.6460 / 8.51 / 0.2120 for He et al. (2013a).

  • Task competition is real and measurable. The paper reports that learning multiple tasks degrades performance on any individual task, motivating the SMoEs design.

  • Learning-strategy comparison (Table 2). Mixed learning gives 97.223 ShapeACC (T1), PSNR 13.07 (T2), 15.74 (T3), and 21.73 (T4) at 357.9M parameters. Sequential-learning orders vary widely: the SL (3-2-4-1) order gives PSNR 17.54 on T2 and 23.55 on T3 but only 15.10 on T4, showing that whichever task is trained first tends to dominate. UniRect (SMoEs) reports 97.390, 19.90, 25.07, and 23.16 with 369.6M parameters. These mixed-learning and sequential-learning runs used a longer 250 epochs.

  • Prompts matter, and matter unevenly (Table 3). Removing prompts drops T1 to 96.324 ShapeACC (from 97.454) and T3 to 23.82 PSNR (from 25.10), has a slight effect on T2 (19.84 versus 19.90), and no effect on T4 (23.15 versus 23.16).

  • RP-TPS ablation (Table 4, labeled T5 in the paper). Without RP-TPS the results are PSNR 21.56, SSIM 0.6502, FID 7.21, LPIPS 0.0972; with RP-TPS they are 23.16, 0.7179, 6.55, 0.0873. The authors note that repeating the grid generator and control-point predictor recursively gave no obvious improvement and added parameters because of the matrix inversion in the recursive form.

  • Mamba versus CNN and Transformer (Table 5). On T4, Mamba gives PSNR 25.10, SSIM 0.7526, FID 19.59, LPIPS 0.1120 at 357.9M parameters, versus CNN at 24.12 / 0.7123 / 21.03 / 0.1250 with 508.9M parameters and Transformer at 24.23 / 0.7226 / 21.27 / 0.1320 with 1.168G parameters. On the task labeled T5, CNN scores 22.31 / 0.6796 / 7.20 / 0.0972, Transformer 23.23 / 0.6884 / 8.59 / 0.1546, and Mamba 23.16 / 0.7179 / 6.55 / 0.0873.

  • Restoration Module contribution is task-dependent (Table 6). On T4, adding the RM improves PSNR from 22.84 to 23.16, SSIM from 0.6901 to 0.7179, FID from 7.88 to 6.55, and LPIPS from 0.1560 to 0.0873. On T2, adding the RM gives PSNR 19.90 versus 19.99 without it but improves FID from 28.69 to 12.51, so in that task the deformation module dominates.

  • Efficiency: 357.9M parameters, 62.98G FLOPs, and 35.8 FPS are reported for computation complexity.

  • Caveat reported by the authors about T1: because the TPS model is irreversible, precise coordinates of key points on the distorted image after rectification are hard to obtain, so the authors use an estimation method that they state substantially reduces the metrics for their method on that task.

Methodology in Plain English

The authors start by asking what the "undo" of each task looks like as a pixel-motion field. Using optical flow visualizations (generated with RAFT), they observe that portrait distortion, wide-angle rectangling, stitched-image rectangling, and rotation correction all produce radial-looking flows around a principal point, differing mainly by which parameters are active. They write one general equation that reduces to each task by zeroing out specific terms, such as the rotation angle, the eccentric translation, or one of the two sets of distortion coefficients.

UniRect then predicts the warp rather than solving the equation directly. A control-point predictor network takes the input image concatenated with a visual prompt and outputs offsets for a grid of control points (12 × 10 in the experiments). Those points define a TPS warp, which produces a sampling grid; a differentiable sampler reads pixels from the original input. The process is repeated once more, with the second predictor refining the points from the first stage, giving the "residual progressive" behavior. Prompts differ by task: a face mask for portrait correction, boundary indicators for the two rectangling tasks, and an all-white image for rotation correction.

A Mamba block sits inside the control-point predictor to scan long-range geometric structure, and the restoration module uses four Residual Mamba Blocks (32 channels in their convolution layers) to repair sampling artifacts. A partial convolution layer handles irregular borders left after warping; for tasks without boundary changes it reduces to a standard convolution. Training combines an appearance (L1) loss, a boundary loss that pulls outermost control points toward the mask boundary, a line/shape penalty on the mesh, and a gradient loss, with the latter three only used where boundary changes occur.

For the joint model, a ResNet18-based gating network scores five expert copies of UniRect and applies a top-1 softmax weighting, so each input is effectively routed to one expert. The idea is that each expert can specialize rather than all tasks competing inside one set of weights.

Training used PyTorch on four NVIDIA Tesla V100 GPUs. Single-task runs took one to three days with batch size 4 and 200 epochs; the joint multi-task run took about 7 days with batch size 10 and 200 epochs. The learning rate followed a poly policy with factor 0.96, starting at 1e-5 for rotation correction and 1e-4 for the remaining tasks, with the Adam optimizer and weight decay 1e-5. Loss weights γ, α1, α2, α3 were set to 0.9, 1e-2, 1.0, and 1e-2. The paper does not report the number of images in each dataset.

Why This Matters

Impact on research. The paper argues that image correction and rectangling should not be treated as separate inverse problems with separate networks, and demonstrates that a shared architecture plus prompts can match or beat task-specific baselines on three of the four tasks studied. It also contributes evidence about how strongly unrelated low-level tasks interfere during joint training, and shows one mechanism (SMoEs with top-1 gating) that largely recovers single-task performance.

Real-world applications:

  • Smartphone camera pipelines, where portrait photos from wide-angle lenses need face-region correction.
  • Rectangling of wide-angle landscape photos that users want as rectangular rather than pincushion-shaped images.
  • Rectangling of stitched panoramas, where the stitched result has irregular borders that simple cropping would ruin.
  • Automatic rotation correction of slightly tilted photographs in galleries or photo-editing apps.
  • Prompt-controlled correction, where the same input image can be corrected differently depending on the prompt, as illustrated with an image containing two distortion types.

Industry relevance. The motivation is explicitly edge deployment: one task-agnostic model instead of four would save memory and computation on devices with limited resources, and 357.9M parameters, 62.98G FLOPs, and 35.8 FPS are reported for that reason. The code is released at https://github.com/yyywxk/UniRect.

Future Directions

  • Simultaneous multiple rectifications. The conclusion raises the scenario of applying several rectifications to the same sample at once, and notes that verification is currently difficult because all existing datasets contain only one type of distortion.
  • Reducing the small joint-training gap. SMoEs is described as "nearly close" to single-task learning rather than equal to it, and the SMoEs model is slightly larger (369.6M versus 357.9M parameters), leaving room for better routing or compression.
  • Pushing perceptual quality on T2. FID and LPIPS on rectified wide-angle image rectangling remain worse than RecRecNet in the reported table, so perceptual fidelity there is an open problem.
  • Deployment-oriented work. The paper states that sharing one network makes it more convenient to design model compression and inference-acceleration algorithms for limited-compute devices, which is presented as follow-up work rather than solved here.

Target Audience

Researchers and graduate students working on low-level vision, image restoration, and computational photography, particularly those interested in multi-task learning, Mamba-based backbones, and warping-based geometric correction. Practitioners building smartphone camera or photo-editing pipelines will benefit from the unified-architecture argument and the efficiency numbers, though the mathematical distortion model and TPS formulation require a solid background to follow. Readers looking for dataset sizes or a fully independent evaluation should note that several implementation details are deferred to supplementary files that are not included in the main content.

Paper metadata: arXiv:2512.18718 [cs.CV]. Authors: Linwei Qiu, Gongzhe Li, Xiaozhe Zhang, Qilin Sun, Fengying Xie (corresponding author). Affiliations include Tianmushan Laboratory and Beihang University, the State Key Laboratory of High-Efficiency Reusable Aerospace Transportation Technology, Huazhong University of Science and Technology, and The Chinese University of Hong Kong, Shenzhen. Funding: National Natural Science Foundation of China Grants 62475006 and 62125102, and National Key Research and Development Program of China Grant 2022ZD0160401.

Authors’ abstract

Image correction and rectangling are valuable tasks in practical photography systems such as smartphones. Recent remarkable advancements in deep learning have undeniably brought about substantial performance improvements in these fields. Nevertheless, existing methods mainly rely on task-specific architectures. This significantly restricts their generalization ability and effective application across a wide range of different tasks. In this paper, we introduce the Unified Rectification Framework (UniRect), a comprehensive approach that addresses these practical tasks from a consistent distortion rectification perspective. Our approach incorporates various task-specific inverse problems into a general distortion model by simulating different types of lenses. To handle diverse distortions, UniRect adopts one task-agnostic rectification framework with a dual-component structure: a {Deformation Module}, which utilizes a novel Residual Progressive Thin-Plate Spline (RP-TPS) model to address complex geometric deformations, and a subsequent Restoration Module, which employs Residual Mamba Blocks (RMBs) to counteract the degradation caused by the deformation process and enhance the fidelity of the output image. Moreover, a Sparse Mixture-of-Experts (SMoEs) structure is designed to circumvent heavy task competition in multi-task learning due to varying distortions. Extensive experiments demonstrate that our models have achieved state-of-the-art performance compared with other up-to-date methods.

Read the original paper