Research
X2HDR: HDR Image Generation in a Perceptually Uniform Space
X2HDR: HDR Image Generation in a Perceptually Uniform Space Overview Research area: Computer vision, specifically high-dynamic-range (HDR) imaging combined with text-to-image diffusion models. Technic

- arXiv
- 2602.04814
- Published
- 2026-02-04
- Authors
- Ronghuan Wu, Wanchao Su, Kede Ma, Jing Liao, Rafał K. Mantiuk
AI summary
X2HDR: HDR Image Generation in a Perceptually Uniform SpaceOverview
Research area: Computer vision, specifically high-dynamic-range (HDR) imaging combined with text-to-image diffusion models.
Technical level: Intermediate. The reader needs some familiarity with latent diffusion models, variational autoencoders (VAEs), and basic HDR concepts such as linear light and tone mapping.
Scope: The paper proposes a single adaptation strategy, X2HDR, that lets existing low-dynamic-range (LDR) pretrained text-to-image diffusion models produce HDR images for two tasks — text-to-HDR synthesis and single-image RAW-to-HDR reconstruction — by operating in a perceptually uniform encoding space instead of linear RGB.
What This Paper Is About
State-of-the-art image generators such as Stable Diffusion and FLUX output LDR images because they are trained on billions of display-encoded, nonlinearly compressed pictures, while HDR and RAW data are natively linear-light and have very different pixel statistics. Prior HDR adaptations work around this by emulating the classic "bracket-and-merge" HDR pipeline, which adds architectural complexity, raises inference cost, and complicates porting to newer backbones. X2HDR's goal is to close that representational gap with a simple preprocessing step — mapping HDR into a perceptually uniform space such as PU21 or PQ — so that an off-the-shelf LDR-pretrained model can handle HDR with only a small amount of finetuning.
Key Contributions
-
A simple HDR representation practice. Encoding HDR and RAW inputs in a perceptually uniform space (PU21 or PQ) allows a frozen, LDR-pretrained VAE to reconstruct HDR content with fidelity comparable to LDR data, avoiding the need for complex bracket-and-merge pipelines or a redesigned VAE decoder.
-
A unified computational method. X2HDR supports both text-to-HDR generation and single-image RAW-to-HDR reconstruction using the same frozen-VAE plus LoRA-finetuned-denoiser recipe, with each task trained under a task-specific LoRA.
-
Empirical evidence that the representation, not finetuning volume, is the bottleneck. Ablations show that linear-space HDR finetuning collapses dynamic range and quality, while PQ and PU21 encodings perform comparably and well.
-
Perceptual validation on a calibrated HDR display. A pairwise comparison study with 26 participants confirms the objective metric trends for both tasks.
Main Findings
-
A frozen LDR VAE can reconstruct PU21-encoded HDR well. On 512 aligned (LDR, HDR) frame pairs cropped and resized to 768×768, the JOD score dropped only from 9.86 for LDR to 9.44 for PU21-encoded HDR, while linear HDR fell to 8.54. Supporting metrics followed the same pattern: exposure-optimized PSNR 38.8 (LDR), 35.6 (PU21 HDR), 27.9 (linear HDR); SSIM 0.961, 0.965, 0.867; LPIPS 0.026, 0.020, 0.129; DISTS 0.008, 0.010, 0.067.
-
Text-to-HDR quality and dynamic range improve over prior methods. On 100 prompts generated by ChatGPT-5 at 512×512 output resolution, the FLUX-based X2HDR reached image quality 0.668 and text-image alignment 0.831 (Q-Eval-100K scores), versus 0.488/0.623 for LEDiff and 0.524/0.671 for Bracket Diffusion. Its effective dynamic range was about 14 stops, compared with 4.3 stops for LEDiff and 11.1 stops for Bracket Diffusion.
-
X2HDR is far cheaper at inference than Bracket Diffusion. The SD-1.5 variant ran in 2.2 seconds with 2.9 GB peak memory; the FLUX variant ran in 11.2 seconds with 33.0 GB. Bracket Diffusion required 557.5 seconds and 38.0 GB, while LEDiff used 7.4 seconds and 8.6 GB.
-
RAW-to-HDR reconstruction is the strongest across all reported metrics. On 96 RAW images from the SI-HDR dataset at 512×512, X2HDR scored JOD 7.57, exposure-optimized PSNR 22.5, SSIM 0.741, LPIPS 0.174 and DISTS 0.089, ahead of RawHDR (6.58, 19.3, 0.591, 0.282, 0.133), LEDiff (6.74, 19.7, 0.570, 0.319, 0.151) and Bracket Diffusion (7.08, 21.6, 0.727, 0.282, 0.148).
-
Perceptually uniform encoding is the critical ingredient. With linear encoding and finetuning, text-to-HDR collapsed to 5.5 stops of dynamic range (image quality 0.578, alignment 0.744). PQ reached 14.5 stops (0.661, 0.822) and PU21 reached 14.2 stops (0.668, 0.831). In RAW-to-HDR, linear encoding gave JOD 5.60 versus 7.48 for PQ and 7.57 for PU21.
-
Finetuning is also necessary. Decoding and rescaling pretrained T2I latents directly, without any finetuning, over-stretched tonal and chromatic range, producing exaggerated contrast and saturation plus widespread clipping.
-
Human observers preferred X2HDR. In the perceptual study, the FLUX-based X2HDR scored 1.81 JOD for text-to-HDR and the SD-based variant scored 0.90 JOD. For RAW-to-HDR, X2HDR scored 1.87 JOD, nearly matching the ground-truth HDR reference, while Bracket Diffusion scored −0.60 JOD and LEDiff and RawHDR showed only statistically insignificant improvements over the raw input.
-
Failure modes of the baselines. Bracket Diffusion can collapse large regions toward near-zero values (a "shadow clipping" mode), while LEDiff often underestimates peak luminance and washes out shadows. Under image conditioning, Bracket Diffusion is also limited to a 256×256 operating resolution.
Methodology in Plain English
The authors start from the observation that HDR and RAW data are stored in linear light, where equal numerical steps do not correspond to equal perceptual steps — human vision is far more sensitive to luminance changes in shadows than in highlights. Under PQ, for example, a 1 cd/m² increase near darkness (0.005 cd/m²) is over 150× more salient than the same increase at 100 cd/m². Because pretrained diffusion models were trained on display-encoded LDR images that already have this kind of perceptual shaping, feeding them linear HDR values creates a mismatch.
X2HDR's fix is to remap HDR values before they reach the network. Linear RGB is globally rescaled so the maximum luminance corresponds to a peak of 4,000 cd/m² (matching commercially available HDR displays), then passed through PU21, a log-quadratic function fitted with parameters a = 0.001908 and b = 0.0078 and offset L_min = log₂(0.005). PU21 compresses extreme highlights and reallocates precision toward shadows, making the statistics closer to sRGB.
With this encoding in place, the VAE is frozen and only the denoiser is adapted, using low-rank adaptation (LoRA) inserted into the query, key, value and output projections of the attention blocks. Training uses the standard flow-matching objective on FLUX.1-dev (and also SD-1.5 as an alternative backbone).
For text-to-HDR, the model maps a text prompt to a PU21-encoded latent and inverts back to linear HDR at the end. For RAW-to-HDR, a RAW capture is demosaicked, rescaled to the same 4,000 cd/m² peak and PU21-encoded; its image tokens are simply concatenated with the text and latent tokens, which works because the FLUX backbone is a diffusion Transformer (DiT) that accepts variable-length token sequences. Both branches are trained separately with task-specific LoRAs but share the same denoising and decoding pipeline.
Evaluation combines Q-Eval-100K (two finetuned vision-language models based on Qwen2-VL-7B-Instruct) for quality and alignment, a percentile-based effective dynamic range measure in stops, and exposure-optimized LDR metrics (PSNR, SSIM, LPIPS, DISTS) alongside ColorVideoVDP JOD scores.
Why This Matters
Impact on research. The paper reframes HDR adaptation as primarily a representation problem rather than an architecture problem. If a frozen LDR-pretrained VAE can already encode HDR faithfully once it is in PU21 space, then much of the recent machinery — separate highlight and shadow denoisers, jointly denoised exposure stacks, modified VAE decoders — may be unnecessary. That lowers the barrier to porting HDR capability onto each new generation of image generator.
Real-world applications:
- Content creation: generating HDR stills directly from a text prompt for film, game, and advertising pipelines, visualized with exposure-adjusted views at EV −4 and EV +4.
- Mobile and camera photography: reconstructing HDR from a single RAW capture, denoising dark regions and inpainting saturated areas instead of relying on multi-exposure bracketing that suffers from ghosting and misalignment.
- ISP integration: since the method targets RAW input before nonlinear in-camera processing, it is positioned as a component inside the camera image signal processor rather than a post-processing step on legacy LDR files.
- HDR display content: producing outputs with roughly 14 stops of effective dynamic range for viewing on HDR monitors.
Industry relevance. The inference cost profile matters for deployment. The SD-1.5 variant of X2HDR runs in 2.2 seconds with 2.9 GB peak memory, versus 557.5 seconds and 38.0 GB for Bracket Diffusion — a difference of more than two orders of magnitude in time, attributable to LoRA finetuning and the absence of bracket generation and merging. Video-game and streaming pipelines, camera firmware teams, and display manufacturers all have a stake in cheap, backbone-agnostic HDR synthesis.
Future Directions
-
Out-of-domain robustness. The current training data are dominated by natural photographs, and the authors report degraded generalization to out-of-domain styles such as cartoons.
-
Extreme exposure handling. For RAW-to-HDR, X2HDR can still fail in extremely underexposed or overexposed regions, occasionally producing implausible hallucinations or local detail inconsistencies.
-
Display-aware and controllable generation. X2HDR is currently display-agnostic: it does not condition on target peak luminance or other device characteristics, and it offers only limited support for controllable HDR generation and interactive HDR editing.
-
Extension to higher resolutions and broader backbones. The headline text-to-HDR comparison was run at 512×512 to match baseline constraints even though X2HDR supports higher resolutions; the authors report additional high-resolution outputs in the supplement, and the LoRA-based design is explicitly motivated by compatibility with newer, more memory-intensive backbones.
Target Audience
Researchers and practitioners in computational photography, HDR imaging, and generative modeling who want to extend diffusion-based image generation beyond LDR output. It is most useful to readers already comfortable with latent diffusion and VAE latent spaces, and to engineers who need a practical, low-overhead path to HDR generation or RAW-domain reconstruction rather than a new architecture. Readers interested in perceptual image quality metrics (JOD, ColorVideoVDP, exposure-optimized PSNR/SSIM) and in human perceptual study design will also find the evaluation protocol relevant.
Authors’ abstract
High-dynamic-range (HDR) formats and displays are becoming increasingly prevalent, yet state-of-the-art image generators (e.g., Stable Diffusion and FLUX) typically remain limited to low-dynamic-range (LDR) output due to the lack of large-scale HDR training data. In this work, we show that existing pretrained diffusion models can be easily adapted to HDR generation without retraining from scratch. A key challenge is that HDR images are natively represented in linear RGB, whose intensity and color statistics differ substantially from those of sRGB-encoded LDR images. This gap, however, can be effectively bridged by converting HDR inputs into perceptually uniform encodings (e.g., using PU21 or PQ). Empirically, we find that LDR-pretrained variational autoencoders (VAEs) reconstruct PU21-encoded HDR inputs with fidelity comparable to LDR data, whereas linear RGB inputs cause severe degradations. Motivated by this finding, we describe an efficient adaptation strategy that freezes the VAE and finetunes only the denoiser via low-rank adaptation in a perceptually uniform space. This results in a unified computational method that supports both text-to-HDR synthesis and single-image RAW-to-HDR reconstruction. Experiments demonstrate that our perceptually encoded adaptation consistently improves perceptual fidelity, text-image alignment, and effective dynamic range, relative to previous techniques.