Skip to content
AI.info

Research

PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation

PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation Overview Research area: Computer vision — polarization imaging, polarimetric representation learning, a

PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation
arXiv
2610.08346
Published
2026-10-06
Authors
Beibei Lin, Tingting Chen, Xin Zhang, Wenhao Zhao, Dongjun Li, Zifeng Yuan

AI summary

PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation

Overview

Research area: Computer vision — polarization imaging, polarimetric representation learning, and benchmark design for physically grounded regression.

Technical level: Intermediate. The core idea is accessible, but the paper assumes familiarity with Stokes vectors, polarization descriptors (AoLP, DoLP, DoCP), and standard image-restoration/generative backbones (Restormer, Uformer, MAE, DiT, WDiff, RealFill, I2ITurbo).

Scope: The paper defines, releases, and empirically analyzes a benchmark that makes the per-scene radiometric scale an explicit prediction and evaluation target for RGB-to-Stokes estimation, rather than something divided out of the problem.

What This Paper Is About

Existing methods that infer polarization from RGB-like images predict only normalized Stokes components or relative descriptors such as AoLP and DoLP. In doing so they divide out the radiometric scale — the per-scene absolute radiance level — which means their outputs cannot be turned into absolute Stokes vectors that radiance-level applications require. PolarScale reformulates the task so that this scale is predicted and scored explicitly, and because the scale is removed from the input by construction, it evaluates dataset-conditioned semantic scale estimation against a constant-scale control rather than claiming the scale is physically recoverable.

Key Contributions

  1. Task reformulation. PolarScale makes the per-scene radiometric scale σ an explicit prediction and evaluation target of RGB-to-Stokes estimation, states the input modality explicitly (the normalized total-intensity image s₀, a scene-referred linear image, not a consumer sRGB photograph), and states the identifiability scope of the problem.

  2. A physics-aware evaluation protocol. Scale errors are scored against a constant-scale control (the training median, 4.85), complemented by wrap-aware AoLP angular error (period 180°), Stokes self-consistency between directly predicted descriptors and those recomputed from predicted Stokes components, and a physical-bound violation rate (fraction of pixels where √(s₁²+s₂²+s₃²) > s₀).

  3. A benchmark and design guidance. Seven restoration-based and generative backbones are evaluated under three prediction strategies — direct Stokes estimation, joint estimation, and decoupled multi-decoder estimation — showing that explicit descriptor supervision and decoupled decoding help and characterizing failure modes such as scale collapse.

  4. Downstream validation with attribution controls. Predicted full-Stokes representations are probed on diffuse/specular separation, material segmentation, and glare classification, with scale-attribution controls that separate the contribution of the learned scale from that of the normalized polarization structure.

Main Findings

  • The strongest restoration models beat the dataset prior on scale. MAE and Uformer reach 3.6% and 4.3% mean relative scale error versus 5.7% for the constant-median control, with Restormer (5.1%) and I2ITurbo (5.3%) beating it more narrowly.

  • Several generative baselines collapse the scale. WDiff reports 92.8% and DiT 99.9% mean relative scale error, with non-positive σ̂ in 11 and 97 of the 200 test scenes respectively; RealFill instead over-estimates the scale at 38.0%.

  • Physical consistency separates restoration from generative models. Restoration models violate the physical bound on fewer than 0.25% of pixels (MAE 0.18%, Uformer 0.22%, Restormer 0.22%), whereas generative models violate it on 3.6–58% of pixels (WDiff 58.1%, DiT 5.18%, RealFill 3.73%, I2ITurbo 3.56%).

  • Explicit descriptor supervision improves descriptor accuracy. Moving from direct estimation to joint or decoupled estimation raises descriptor PSNR for every backbone — for MAE from 18.88 dB to 23.54/23.66 dB, and for I2ITurbo from 16.48 dB to 21.77/21.49 dB. Decoupled estimation avoids the trade-off against scale-dependent outputs that joint estimation can introduce (I2ITurbo's normalized-Stokes PSNR drops from 38.89 dB to 34.98 dB under joint estimation).

  • Full-Stokes PSNR is a poor measure of scale recovery. Replacing MAE's learned σ̂ with the training median changes its full-Stokes PSNR from 21.08 dB to 21.13 dB, and the training mean, whose scale error is roughly three times larger (16.2%), yields an even higher PSNR for every model. WDiff's full-Stokes PSNR falls from 18.30 dB to 2.65 dB once an honest constant scale is applied, exposing that its inflated value came from a collapsed reconstruction.

  • The learned scale is not merely a re-learned constant, but it regresses toward the prior. MAE's deviation of σ̂ from the training prior correlates with the ground-truth deviation at r = 0.43, yet on scenes whose σ departs strongly from the prior the models regress toward it.

  • Descriptor accuracy remains limited. Even with direct supervision, the best AoLP angular error across evaluated models is 26.1°.

  • Downstream tasks improve with predicted full-Stokes representations. Diffuse/specular AbsRel drops from 0.190/0.508 (RGB-direct regression) to 0.118/0.445; material segmentation mIoU rises from 0.330 with predicted DoLP/AoLP to 0.374 with predicted ŝₙ plus σ̂; glare accuracy over the 200 test scenes rises from 48.5% (Wilson 95% interval [41.7, 55.4]) to 63.0% ([56.1, 69.4]).

  • The contribution of the learned scale itself is mixed. In diffuse/specular separation the learned σ̂ performs on par with the constant control (0.118/0.445 versus 0.118/0.448), while the ground-truth scale gives 0.100/0.440 and a shuffled scale degrades the diffuse estimate to 0.133. In material segmentation, adding σ̂ to predicted normalized Stokes raises mIoU from 0.355 to 0.374, but the probe lacks constant- and shuffled-scale variants to isolate the per-scene scale.

  • Native multi-channel heads were the weaker adaptation for prompt-steered generative models. Retraining RealFill and I2ITurbo with native heads raised I2ITurbo's scale error from 5.3% to 8.8%, collapsed RealFill's scale (σ̂ ≤ 0 in 21 of 200 scenes, r = −0.07 with the true scale), and improved only self-consistency via single-pass prediction (I2ITurbo 27.4° versus 52.4°).

Methodology in Plain English

The researchers did not capture new data. They built the benchmark on the existing trichromatic full-Stokes measurements of the Spectral-Polarization dataset (2,022 trichromatic and 311 hyperspectral images; trichromatic images at 2100 × 1920), which retain the radiometric scale.

For each scene they define a per-scene scale σ as the maximum of S₀ over the scene, then divide every Stokes component by σ to get normalized components s₀–s₃. The network input is the normalized total-intensity image s₀, and the model must predict the normalized Stokes components s₁–s₃, the descriptors AoLP/DoLP/DoCP, and the per-scene scale σ, from which full Stokes components follow as Sₙ = σ · sₙ. The output is a 19-channel polarimetric representation (9 normalized Stokes channels, 1 scale map, 9 descriptor channels), with 10 channels for the direct strategy.

They adopt the 1,000/200 train/test scene split of Lin et al. so results are directly comparable to that protocol. Because σ is divided out of the input, the paper is explicit that it cannot be recovered by physical inversion — a globally rescaled scene yields the same input — so it is treated as a semantic estimation problem conditioned on the dataset, analogous to metric monocular depth estimation.

Seven backbones are trained to convergence on the same server with eight NVIDIA RTX A5000 GPUs (24 GB each), taking about one day per restoration model and two days per diffusion-based model. Three prediction strategies are compared: predicting σ and normalized Stokes directly and deriving descriptors, predicting everything with a unified decoder, and using two parallel decoders for scale-dependent versus scale-independent quantities. Evaluation pairs standard image-quality metrics (PSNR with MAX_I = 1.0, SSIM, and LPIPS, which is scale-invariant and reported once) with the physics-aware metrics and a constant-σ control.

Why This Matters

Impact on research. PolarScale reframes a task that prior work had implicitly defined around scale-free outputs, and it supplies a protocol, preprocessed splits, scale values, derived targets, evaluation scripts, and a datasheet in a standalone package. It also demonstrates a benchmarking pattern — scoring an underdetermined target against a trivial control rather than treating a single aggregate score as evidence — that extends beyond polarization.

Real-world applications:

  • Diffuse/specular separation, where the polarized part of reflected radiance is attributed to the specular component in absolute radiance terms.
  • Glare-level assessment, since glare depends on the amount of polarized light reaching the sensor, not on the fraction of light that is polarized.
  • Material segmentation and recognition, via polarization cues that vary with surface roughness and coating.
  • Robotics, autonomous systems, and wearable sensing, where polarization-aware perception could reduce reliance on dedicated polarimetric hardware.

Industry relevance. The work targets settings where polarization hardware is costly, reduces spatial resolution through multiplexed designs, and complicates calibration — mobile, wearable, and robotics platforms. The paper is explicit that predicted scales are dataset-conditioned estimates rather than calibrated radiometry and warns against over-reliance in safety-critical settings such as driver assistance before validation under the target sensor, lighting, and domain.

Future Directions

  • Improving scale accuracy beyond the prior. The learned scale does not consistently beat the constant control downstream, and models regress toward the training prior on scenes whose scale departs strongly from it; scale accuracy is identified as the open problem PolarScale is designed to track.

  • Broadening data coverage. All data come from one dataset captured with one polarimetric camera, and the scale distribution is narrow (mean 4.25, median 4.85 in training), leaving only a small margin for learned models. Domain generalization and test-time adaptation are suggested, and scale results may not transfer to other sensors, exposures, or scene distributions.

  • Extending to consumer RGB. The current input is a normalized linear image captured through polarimetric optics; consumer sRGB adds ISP nonlinearities (tone mapping, white balance, clipping) and polarization-dependent sensor responses. No dataset pairs consumer RGB with absolute-scale Stokes ground truth, so this setting cannot be evaluated quantitatively yet.

  • Handling harder scene conditions and fine structure. Low light, strong specular highlights, and adverse weather are not covered, and all models still fail on fine structures and abrupt polarization transitions, which physics-informed constraints may address.

Target Audience

Researchers and engineers working on polarization imaging, computational photography, and physically grounded inverse problems in vision; benchmark designers interested in evaluating targets that are underdetermined by their inputs; and practitioners in robotics, material analysis, remote sensing, and computational imaging who want to know whether absolute Stokes information can be inferred from intensity images and where current methods stand.

Authors’ abstract

Polarization imaging provides physical cues beyond intensity imaging but typically requires specialized hardware. Recent methods infer polarization from RGB-like inputs, yet predict only normalized Stokes components or relative descriptors, from which the radiometric scale needed for full Stokes reconstruction has been divided out. We introduce PolarScale, a benchmark that makes this scale an explicit prediction and evaluation target. Built on existing trichromatic full-Stokes measurements, PolarScale takes the per-scene normalized total-intensity image $s_0$ (a scene-referred linear image, not a consumer sRGB photograph) and asks models to predict normalized Stokes components, AoLP/DoLP/DoCP, and a per-scene scale. Because the scale is divided out of the input, it is not physically identifiable; PolarScale therefore evaluates dataset-conditioned semantic scale estimation against a constant-scale control, together with angular, self-consistency, and physical-bound metrics. Across seven restoration-based and generative backbones and three prediction strategies, the strongest restoration models estimate the scale with 3.6-4.3% mean relative error versus 5.7% for the constant control and violate physical bounds on fewer than 0.25% of pixels, whereas two generative baselines collapse to a near-zero scale; explicit descriptor supervision improves descriptor accuracy (23.66 vs. 18.88 dB PSNR for MAE). Predicted full-Stokes representations improve diffuse/specular separation, material segmentation, and glare classification, although in diffuse/specular separation the learned scale performs only on par with the constant control.

Read the original paper