Skip to content
AI.info

Research

NTIRE 2025 Challenge on Low Light Image Enhancement: Methods and Results

Overview Research area: Computer vision — low-light image enhancement (LLIE), reported as a challenge summary from the NTIRE 2025 Workshop. Technical level: Intermediate. The competition structure and

arXiv
2510.13670
Published
2025-10-15
Authors
Xiaoning Liu, Zongwei Wu, Florin-Alexandru Vasluianu, Hailong Yan, Bin Ren, Yulun Zhang, Shuhang Gu, Le Zhang, Ce Zhu, Radu Timofte, Kangbiao Shi, Yixu Feng, Tao Hu, Yu Cao, Peng Wu, Yijin Liang, Yanning Zhang, Qingsen Yan, Han Zhou, Wei Dong, Yan Min, Mohab Kishawy, Jun Chen, Pengpeng Yu, Anjin Park, Seung-Soo Lee, Young-Joon Park, Zixiao Hu, Junyv Liu, Huilin Zhang, Jun Zhang, Fei Wan, Bingxin Xu, Hongzhe Liu, Cheng Xu, Weiguo Pan, Songyin Dai, Xunpeng Yi, Qinglong Yan, Yibing Zhang, Jiayi Ma, Changhui Hu, Kerui Hu, Donghang Jing, Tiesheng Chen, Zhi Jin, Hongjun Wu, Biao Huang, Haitao Ling, Jiahao Wu, Dandan Zhan, G Gyaneshwar Rao, Vijayalaxmi Ashok Aralikatti, Nikhil Akalwadi, Ramesh Ashok Tabib, Uma Mudenagudi, Ruirui Lin, Guoxi Huang, Nantheera Anantrasirichai, Qirui Yang, Alexandru Brateanu, Ciprian Orhei, Cosmin Ancuti, Daniel Feijoo, Juan C. Benito, Álvaro García, Marcos V. Conde, Yang Qin, Raul Balmez, Anas M. Ali, Bilel Benjdira, Wadii Boulila, Tianyi Mao, Huan Zheng, Yanyan Wei, Shengeng Tang, Dan Guo, Zhao Zhang, Sabari Nathan, K Uma, A Sasithradevi, B Sathya Bama, S. Mohamed Mansoor Roomi, Ao Li, Xiangtao Zhang, Zhe Liu, Yijie Tang, Jialong Tang, Zhicheng Fu, Gong Chen, Joe Nasti, John Nicholson, Zeyu Xiao, Zhuoyuan Li, Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, Subrahmanyam Murala, Duan Liu, Weile Li, Hangyuan Lu, Rixian Liu, Tengfeng Wang, Jinxing Liang, Chenxin Yu

AI summary

Overview

Research area: Computer vision — low-light image enhancement (LLIE), reported as a challenge summary from the NTIRE 2025 Workshop.

Technical level: Intermediate. The competition structure and metrics are easy to follow; the individual team methods assume familiarity with deep learning architectures (Transformers, CNNs, diffusion, state-space models).

Scope: A review of the NTIRE 2025 Low-Light Image Enhancement Challenge, covering its ranking protocol, the participating teams' methods, and the final leaderboard of 28 valid submissions.

What This Paper Is About

Low-light photographs come out dark, noisy, colour-distorted, and full of artifacts, and correcting them often introduces new problems. This paper reports on a competition that asked teams to build networks that turn such images into brighter, clearer, and more visually compelling results across diverse hard conditions — dim scenes, severe darkness, backlighting, non-uniform illumination, and indoor and outdoor night scenes. The paper summarizes what each team built and how they ranked.

Key Contributions

  1. Ran and documented the NTIRE 2025 LLIE Challenge, the successor to the NTIRE 2024 LLIE Challenge, drawing 762 registered participants of whom 28 teams submitted valid entries.
  2. Defined a composite ranking metric combining PSNR, SSIM, LPIPS, and NIQE with weights of 0.5, 0.5, 0.4, and 0.2 respectively, reported alongside per-metric ranks.
  3. Compiled descriptions of the participating methods, from ESDNet and Retinexformer variants to Transformer, Mamba, diffusion, normalizing-flow, and Schrödinger Bridge approaches, plus their training configurations.
  4. Enforced reproducibility, requiring top-performing teams to submit training scripts and excluding one team (JHC-Info) that failed to provide a checkpoint within the competition period.

Main Findings

  • Participation: 762 participants registered and 28 teams submitted valid entries.
  • Dataset scale: The dataset contains 219 training scenes, 46 validation scenes, and 30 test scenes, with image resolutions reaching 4K and beyond. Ground truth for validation and testing was hidden from participants. The paper states that detailed dataset specifications will be published in future work.
  • Winner: NWPU-HVI took Final Rank 1 with PSNR 26.24, SSIM 0.861, LPIPS 0.128, and NIQE 10.95, using a three-model fusion (ESDNet, Retinexformer, CIDNet) combined by a linear fusion scheme.
  • Runner-up: Imagine ranked 2nd with the top PSNR of 26.35, plus SSIM 0.858, LPIPS 0.133, and NIQE 11.81, using a multi-scale CNN-Transformer hybrid UNet called SG-LLIE.
  • Third place: pengpeng-yu ranked 3rd with PSNR 25.85, SSIM 0.858, LPIPS 0.134, and NIQE 11.29, extending the EDSNet implementation from SYSU-FVL-T2's NTIRE 2024 entry.
  • Strongest single-metric scores: Imagine achieved the best PSNR (26.35); DAVIS-K achieved the best SSIM (0.863); MRT-LLIE and SynLLIE tied for the best LPIPS at 0.117; ImageLab achieved the best NIQE at 9.68.
  • Weakest entry: CV-SVNIT ranked 27th with PSNR 16.85, SSIM 0.565, LPIPS 0.427, and NIQE 12.29 — far below the rest of the field.
  • Excluded entries: JHC-Info is excluded from the ranking for failing to supply a reproducible checkpoint. Smartdsp, BUPTMM, SynLLIE, and CV-SVNIT are marked as excluded from the report, and the paper notes that some teams provided fact sheets but did not participate in the challenge report.
  • Narrow margins at the top: The top five teams span PSNR 25.14 to 26.35 and SSIM 0.856 to 0.863, indicating a tightly clustered field on the composite metric.
  • Recurring design patterns: ESDNet and its components (Dilated Residual Dense Blocks, Semantic-Aligned Scale-Aware Modules) appear in many leading entries; many teams also used self-ensembling at inference and progressive training with growing patch sizes.
  • Resolution handling: Several teams processed full-resolution images directly but split the largest inputs — for example, NJUPT-IPR split 4000×6000 images into four tiles, and MRT-LLIE split 4000×6000 images into four 4000×1500 pieces using pixel interleaving.
  • Efficiency focus in one entry: AVC2's MobileIE reports 0.57M parameters for Lux Themps' SADe-ViT, while LR-LL reports that its model converts to TFLite and processes a 2000×2992 image in roughly 100 ms on a Qualcomm Adreno 735 with about 400 MB of memory.

Methodology in Plain English

The organizers reused the setup of the 2024 edition: a hidden-ground-truth image dataset spanning many lighting conditions, split into training, validation, and test scenes. Teams trained on the 219 provided training pairs, submitted enhanced outputs for 46 validation images to a server that returned PSNR and SSIM scores in real time, and then submitted enhanced results plus code and a fact sheet for 30 test images to Codalab.

Because no single metric captures enhancement quality, the organizers ranked teams with a composite score blending PSNR, SSIM, LPIPS, and NIQE, and also reported ranks for each metric individually. They verified results, required training scripts from the top teams for reproducibility, and shortened each team's own method description to fit an eight-page limit.

The methods themselves fall into recognizable families: ESDNet-based encoder-decoders with semantic-aligned multi-scale modules, Retinexformer and Retinex-theory variants, Transformer and hybrid CNN-Transformer networks, Mamba/state-space models, normalizing flows, diffusion-based pipelines, and ensemble combinations of several of these. Training recipes vary widely in patch size, batch size, iterations, and optimizers, with cyclic cosine annealing and self-ensembling appearing frequently.

Why This Matters

Low-light enhancement is a practical bottleneck for photography, surveillance, and any deployment where sensors must work in poor lighting. By running the challenge on a common dataset with hidden ground truth and a fixed composite metric, the organizers made it possible to compare very different architectural families on equal footing and to see which design choices actually move the numbers. The reproducibility requirement — and the exclusion of a team that could not supply a checkpoint — pushes the subfield toward verifiable results rather than unreproducible leaderboard gains.

Real-world applications implied by the entries and setup:

  • Night photography rendering and post-processing of dark smartphone photos, including backlit and non-uniformly lit scenes.
  • On-device camera pipelines, as demonstrated by LR-LL's TFLite conversion, Qualcomm Adreno 735 deployment, and reported 100 ms processing for a 2000×2992 image.
  • Ultra-high-resolution image handling at 4K and beyond, relevant to professional imaging and large-sensor outputs up to 8K in one entry's claim.
  • Video conferencing and video quality enhancement, which is listed among the co-located NTIRE 2025 challenges alongside this one.

Industry relevance: The challenge is sponsored by ByteDance, Meituan, Kuaishou, and the University of Würzburg Computer Vision Lab, and the method descriptions reference production-oriented constraints such as GPU memory limits, tiling strategies for 4000×6000 images, and model sizes measured in parameters — all signals of direct interest to consumer camera and content-platform developers.

Future Directions

  • Publishing the full dataset specification, which the paper explicitly defers to future work, including the precise conditions and resolutions it covers.
  • Closing the reproducibility gap, since one team had to be excluded for a missing checkpoint and one method's pre-trained weights were noted as possibly non-reproducible beyond AMD hardware.
  • Reconciling metric disagreement, as the per-metric ranks diverge sharply from the composite ranking — for example, ImageLab ranked 22nd overall while achieving the best NIQE (9.68), and CV-SVNIT ranked last on PSNR, SSIM, and LPIPS but 26th on NIQE.
  • Improving efficiency at high resolution, given that many entries rely on tiling or downsampling workarounds for 4000×6000 inputs and only a few report lightweight parameter counts or inference latencies.

Target Audience

Researchers and engineers working on image restoration, computational photography, and camera ISP pipelines; participants preparing for or reviewing NTIRE-style challenges; and practitioners who need a current snapshot of which architectures and training recipes perform well on difficult low-light data at high resolution.

Authors’ abstract

This paper presents a comprehensive review of the NTIRE 2025 Low-Light Image Enhancement (LLIE) Challenge, highlighting the proposed solutions and final outcomes. The objective of the challenge is to identify effective networks capable of producing brighter, clearer, and visually compelling images under diverse and challenging conditions. A remarkable total of 762 participants registered for the competition, with 28 teams ultimately submitting valid entries. This paper thoroughly evaluates the state-of-the-art advancements in LLIE, showcasing the significant progress.

Read the original paper