Skip to content
AI.info

Research

A Large-Depth-Range Layer-Based Hologram Dataset for Machine Learning-Based 3D Computer-Generated Holography

Summary: A Large-Depth-Range Layer-Based Hologram Dataset for Machine Learning-Based 3D Computer-Generated Holography Overview Research area: Computer-generated holography (CGH) and machine learning f

A Large-Depth-Range Layer-Based Hologram Dataset for Machine Learning-Based 3D Computer-Generated Holography
arXiv
2512.21040
Published
2025-12-24
Authors
Jaehong Lee, You Chan No, YoungWoo Kim, Duksu Kim

AI summary

Summary: A Large-Depth-Range Layer-Based Hologram Dataset for Machine Learning-Based 3D Computer-Generated Holography

Overview

Research area: Computer-generated holography (CGH) and machine learning for 3D displays — specifically dataset construction and hologram generation quality.

Technical level: Advanced. The paper combines wave-optics simulation (angular spectrum method), layer-based hologram generation, and deep learning benchmarks, and reports optical as well as numerical reconstructions.

Scope: The paper introduces KOREATECH-CGH, a public 6,000-sample RGB-D/hologram dataset with a large depth range, plus an amplitude-projection-enhanced layer-based generation method (AP-LBM) and a focal-image-projection (FIP) evaluation metric.

What This Paper Is About

Machine-learning-based computer-generated holography (ML-CGH) has advanced quickly, but progress is held back because few high-quality, large-scale hologram datasets are publicly available. The best-known public dataset, MIT-CGH-4K, has a narrow 6 mm depth range and an optical configuration tuned to one specific device, while many newer datasets are private and hard to reproduce. This paper builds and releases a large-depth-range, multi-resolution dataset together with a new generation method that keeps reconstructions sharp even when scene depth is large.

Key Contributions

  1. KOREATECH-CGH dataset: 6,000 pairs of RGB-D images and complex holograms, publicly released, spanning resolutions from 256×256 to 2048×2048 and depth ranges up to the theoretical limits of the angular spectrum method.
  2. Amplitude projection (AP-LBM): A post-processing technique that replaces the amplitude components of the hologram wavefield at each depth layer while preserving phase, improving reconstruction fidelity at large depth ranges.
  3. Focal Image Projection (FIP): A quantitative evaluation method, extended from the hologram focal loss, that compares only in-focus regions of a hologram's focal stack against the rendered image, making 3D hologram quality measurable with PSNR and SSIM.
  4. Benchmarking experiments: Training of three hologram generation models (TensorHolography, U-Net, Swin-Unet) and two hologram upscaling models (H2HSR Swin, H2HSR RDN) from scratch on the dataset to demonstrate its usefulness for generation and super-resolution tasks.

Main Findings

  • AP-LBM leads on average and maximum metrics. On 500 holograms at 512×512 resolution, AP-LBM reached an average PSNR_FIP of 23.99 dB and average SSIM_FIP of 0.75, with maximum values of 27.01 dB and 0.87 — versus ADV-LBM's average of 21.96 dB / 0.71 and SM-LBM's average of 21.15 dB / 0.68. The abstract states this surpasses a recent optimized silhouette-masking layer-based method by 2.03 dB and 0.04 SSIM.
  • Minimum scores go slightly to ADV-LBM. ADV-LBM edges out AP-LBM on minimum PSNR (19.77 dB vs 19.71 dB) and minimum SSIM (0.47 vs 0.44); the paper calls the difference marginal and not visually significant.
  • Quality degrades with depth range. Average PSNR_FIP falls from 30.75 dB at 1 mm depth and 27.62 dB at 6 mm to 25.90 dB at 11 mm, 24.74 dB at 16 mm, and 23.99 dB at 20.334 mm. Beyond 11 mm, average metrics drop below 30 dB and 0.9.
  • MIT-CGH-4K's 6 mm range is effectively much smaller in this configuration. The paper notes that, adjusted by the ratio of maximum propagation distance to depth range, MIT-CGH-4K's 6 mm corresponds to roughly 1.6 mm under the KOREATECH-CGH setup (MIT-CGH-4K: 8 µm pixel pitch, 384×384 resolution, 76.98 mm max propagation distance; KOREATECH-CGH: 3.6 µm pitch, 20.72 mm propagation range).
  • Layer count helps up to about 100, then plateaus. Average PSNR_FIP rose from 19.94 dB at 10 layers to 23.11 dB at 100 layers, then to 23.94 dB at 1,000 and 23.99 dB at 10,000. Interestingly, the 100-layer configuration gave the best minimum scores (20.11 dB, SSIM 0.48), suggesting more consistent quality across scenes.
  • Artifact-reduction tricks did not help in the final configuration. Ringing reduction (average 23.83 dB, SSIM 0.73) and edge padding (average 23.19 dB, SSIM 0.73) both scored lower than plain AP-LBM, so neither was applied to the released dataset. However, edge padding did outperform zero-padding when only a small number of layers was used (e.g., 21.22 dB / 0.65 at 10 layers versus 19.94 dB / 0.45).
  • Model complexity correlated with quality on this dataset. In the trained-from-scratch experiments, PSNR for numerical reconstruction improved from 25.3 dB (TensorHolography, 149K parameters) to 28.1 dB (U-Net, 13M) to 28.7 dB (Swin-Unet, 42M). The paper notes that on MIT-CGH-4K, TensorHolography was the strongest model despite having fewer parameters, which the authors attribute to differences in depth range and generation method (layer-based here versus point-based there).
  • Upscaling works but with a cost trade-off. H2HSR Swin (12M parameters) scored higher than H2HSR RDN (22M parameters) on reconstruction PSNR and SSIM (27.2 dB / 0.76 vs 25.7 dB / 0.72), but took 640 ms per sample versus 148 ms, measured on a single NVIDIA RTX A6000 GPU.

Methodology in Plain English

The authors built a synthetic pipeline: they took 100 3D-scanned objects from the Google Scanned Objects (GSO) dataset, randomly resized each to 20–30% of the hologram width, rotated them with Euler angles sampled from [−2π, 2π), and arranged them on a two-dimensional uniform grid so their positions were spread evenly in the image plane. Depth was sampled independently per object from an interval bounded by the angular spectrum method's maximum effective propagation distance (Eq. 1), which is derived from pixel count, pixel pitch, and wavelength. Rendering used OptiX 8.0 with orthographic projection to produce single-precision RGB and depth maps; the red channel wavelength (e.g., 638 nm) determined the shortest valid maximum propagation distance.

Holograms were then generated from these RGB-D pairs using a layer-based approach. The scene is sliced into layers along the depth axis; each layer's wavefront is numerically propagated to the hologram plane using the angular spectrum method, and masking planes ensure only the correct wavefronts contribute. The authors start from silhouette masking (SM-LBM), which masks only backward layers and produces contour aliasing and shadow artifacts. They upgrade it to a bidirectional masking scheme (ADV-LBM) that considers both forward and backward directions and operates over k = 2 layers, suppressing inter-layer contour lines. Finally, amplitude projection (AP-LBM) propagates the finished hologram back to the farthest focal plane, replaces the amplitude there with the corresponding image content while keeping the phase, and repeats this iteratively layer by layer back to the hologram plane. This restores sharp detail that defocus blur from front planes would otherwise erase in the background. The released holograms use 10,000 layers, a fixed 3.6 µm pixel pitch, and 638/532/450 nm wavelengths for red, green, and blue.

For evaluation the authors used FIP, which assembles a composite image from the amplitude of the focal stack using depth-aware masking, so out-of-focus regions are not unfairly penalized. They also built an optical reconstruction test bench using a May HoloKit with an LCoS IRIS-U62 phase modulator (3.6 µm pitch, 3840×2160) and WikiOptics lasers at the same three wavelengths, plus a 4f system, double-phase amplitude encoding, and a 1.1° off-axis angle, imaged with a Nikon 50 mm lens and FLIR Blackfly S camera. For the machine-learning benchmarks, 5,000 samples were used for training, 500 for validation, and 500 for testing.

Why This Matters

This work addresses a concrete bottleneck in ML-CGH: without shared, high-quality, large-depth-range datasets, results are hard to reproduce and compare across labs. By releasing KOREATECH-CGH publicly, the authors give the community a common benchmark with configuration diversity (four resolutions, up to roughly 81 mm depth at the largest resolution) and a documented generation pipeline.

Real-world applications named or implied by the paper:

  • AR/VR holographic displays, where the paper explicitly positions ML-CGH as enabling real-time 3D holography.
  • Next-generation 3D displays that reproduce depth cues and occlusion effects through wave optics.
  • Depth-aware visualization, listed among the future areas the dataset can enable.
  • Real-time holographic rendering pipelines, also named as a target application.
  • Near-eye displays, the scenario for which the earlier MIT-CGH-4K dataset was designed and which motivates accurate hologram generation and upscaling.

Industry relevance comes from the practical framing: real-time hologram generation on consumer hardware (even mobile devices, per the cited Tensor Holography work) is the commercial goal, and a public dataset lowers the barrier for companies and labs to train and compare models on a harder, more realistic configuration than a 6 mm depth range allows.

Future Directions

  • Hybrid generation methods. The authors plan to combine layer-based methods' interpretability and simplicity with the expressiveness of point- and volume-based representations.
  • Dataset expansion. They aim to add dynamic scenes and broader variation in lighting and material properties to better match real holographic display systems.
  • Fixing artifact and blur limitations. AP-LBM still suffers ringing artifacts and unnatural blur in defocused regions caused by smooth phase initialization; random phase initialization could reduce the blur but risks speckle noise in focused regions. The paper cites LDI and Gaussian Splatting as high-quality alternatives that share these problems while requiring much heavier input pre-processing.
  • Parallax and viewing angle. Layer-based methods reproduce parallax poorly at extreme viewing angles, an open problem the authors acknowledge but do not solve.

Target Audience

Researchers and engineers working on computer-generated holography, machine-learning-based display systems, and 3D reconstruction who need a large-depth-range, multi-resolution dataset for training or benchmarking. It is also useful for AR/VR display developers evaluating hologram generation and super-resolution models, and for anyone studying how depth range and layer count trade off against reconstruction quality. Readers need familiarity with wave optics or deep learning pipelines to get the most from the technical sections, though the dataset itself is usable by ML practitioners who only need RGB-D and hologram pairs.

Authors’ abstract

Machine learning-based computer-generated holography (ML-CGH) has advanced rapidly in recent years, yet progress is constrained by the limited availability of high-quality, large-scale hologram datasets. To address this, we present KOREATECH-CGH, a publicly available dataset comprising 6,000 pairs of RGB-D images and complex holograms across resolutions ranging from 256*256 to 2048*2048, with depth ranges extending to the theoretical limits of the angular spectrum method for wide 3D scene coverage. To improve hologram quality at large depth ranges, we introduce amplitude projection, a post-processing technique that replaces amplitude components of hologram wavefields at each depth layer while preserving phase. This approach enhances reconstruction fidelity, achieving 27.01 dB PSNR and 0.87 SSIM, surpassing a recent optimized silhouette-masking layer-based method by 2.03 dB and 0.04 SSIM, respectively. We further validate the utility of KOREATECH-CGH through experiments on hologram generation and super-resolution using state-of-the-art ML models, confirming its applicability for training and evaluating next-generation ML-CGH systems.

Read the original paper