Skip to content
AI.info

Research

An evaluation of SVBRDF Prediction from Generative Image Models for Appearance Modeling of 3D Scenes

Overview Research area: computer graphics and computer vision, specifically inverse rendering, material (SVBRDF) estimation, and generative image models for 3D content creation. Technical level: Inter

An evaluation of SVBRDF Prediction from Generative Image Models for Appearance Modeling of 3D Scenes
arXiv
2512.13950
Published
2025-12-15
Authors
Alban Gauthier, Valentin Deschaintre, Alexandre Lanvin, Fredo Durand, Adrien Bousseau, George Drettakis

AI summary

Overview

Research area: computer graphics and computer vision, specifically inverse rendering, material (SVBRDF) estimation, and generative image models for 3D content creation. Technical level: Intermediate (assumes familiarity with BRDF terminology, diffusion models, and standard image-quality metrics, though the paper's core argument is accessible). Scope: This paper evaluates how different single-image SVBRDF prediction architectures and input choices behave when fed generated (rather than photographed) images of 3D indoor scenes, and builds a fast pipeline that turns untextured scene geometry plus a text or image prompt into a relightable SVBRDF texture atlas.

What This Paper Is About

Texturing a 3D scene with physically based materials normally requires manually authoring albedo, roughness, and metallic maps. Generative diffusion models can already synthesize realistic RGB images of a scene that match its geometry, and single-image SVBRDF predictors can recover material parameters from RGB images, so combining the two suggests a fast route to material atlases. The problem is that a predictor applied independently to each generated view may produce maps that disagree across views, creating seams and blur when merged into a single texture atlas; this paper systematically studies which architectural and input choices minimize that failure.

Key Contributions

  1. An analysis of the design space of single-image SVBRDF estimation for multiview material generation, comparing architectures ([ZLH22], [KSN24], [ZDG24], GenPercept [XGL*25], and a plain UNet) on both per-view accuracy and multiview coherence.
  2. An assessment of alternative input channels to the predictor: geometry buffers (depth and normals) and "hyperfeatures" extracted from the denoising UNet of Stable Diffusion (the UNet-HF design).
  3. A fast and controllable pipeline that generates SVBRDF textures over indoor scenes, combining conditional image generation, reprojection and inpainting with per-view SVBRDF prediction and photogrammetry-style texture merging.
  4. Released code, model weights, and a curated dataset (project page repo-sam.inria.fr/nerphys/svbrdf-evaluation, code at github.com/graphdeco-inria/svbrdf-evaluation).

Main Findings

  • A standard UNet is competitive. The simple UNet-RGB predictor with geometry cues performs surprisingly well on all metrics, being outperformed only by GenPercept (with or without additional geometric information); UNet-HF with geometry cues places third. On the InteriorVerse test set (2633 images), UNet-RGB reaches basecolor PSNR 21.66 / SSIM 0.803 / LPIPS 0.109, roughness 17.73 / 0.672 / 0.241, and metallic 20.13 / 0.843 / 0.256, versus GenPercept's 22.67 / 0.793 / 0.113, 19.26 / 0.689 / 0.202, and 20.83 / 0.847 / 0.254.
  • Accuracy and detail trade off across methods. GenPercept is the most accurate numerically but struggles to recover fine details such as thin structures of a lamp and wall panel. Kocsis et al. and RGB↔X produce smooth, piecewise-constant maps (matching the character of ground-truth SVBRDFs) but these are not always accurate, particularly for roughness and metallic where different objects of the same material receive different, often erroneous values. MGNet tends to produce splotchy results in all maps.
  • Hyperfeatures give the best multiview coherence. UNet-HF achieves the lowest flicker scores in the consistency analysis (basecolor 0.056, roughness 0.050, metallic 0.041), better than UNet-RGB (0.064 / 0.078 / 0.060), GenPercept (0.072 / 0.104 / 0.069), and [ZLH*22] (0.073 / 0.100 / 0.105). A possible interpretation offered is that deep features from the image generator have little dependence on viewpoint, yielding similar material values for objects seen from different views.
  • Generative predictors are too slow for an interactive pipeline. Regression methods need a single inference pass, whereas diffusion models require 50 inferences. Measured on an RTX 3090, UNet-RGB takes 36 ± 1 ms and GenPercept 69 ± 4 ms, while Kocsis et al. takes about 6 s per single prediction and 60.8 ± 1.3 s when 10 predictions are averaged as recommended; Zeng et al. takes 11.1 ± 0.2 s.
  • Averaging reduces detail. Averaging 10 predictions increases coherence for Kocsis et al., but in the authors' experiments this tends to reduce overall contrast and details.
  • Geometry inputs help, but only up to a point. Providing depth and normals as additional channels consistently improves the consistency of estimations and tends to improve overall accuracy; however, the authors conclude that these extra channels provide only marginal results improvements overall.
  • Seams are visible in merged atlases. Kocsis et al. and RGB↔X create seams in the merged texture atlas rendered from a novel viewpoint, while UNet-RGB, UNet-HF, and GenPercept produce more contrasted results with reduced seam artifacts.
  • UNet-HF was chosen for the final results because it achieves the highest multi-view consistency with competitive single-view accuracy; the authors also note UNet-HF shows the lowest level of flickering in fly-through videos.

Methodology in Plain English

The pipeline starts from the untextured geometry of an indoor 3D scene and a set of viewpoints; all results in the paper use five viewpoints, with the first supplied by the user and the other four generated at an offset of 25 degrees in latitude and longitude. For the first view, depth and contour maps are rendered and used to condition ControlNet with a Stable Diffusion backbone to produce a photorealistic, geometry-aligned image; the user can optionally steer content with a text prompt or an example image (encoded via IP-Adapter), and conditions are combined with the Multi-ControlNet pipeline from Diffusers.

Each subsequent view is produced by projecting already-generated images into the new camera using their cameras and depth maps, then inpainting the holes left by disocclusions with a ControlNet trained for inpainting, also conditioned on depth, contours, the text prompt, and a text embedding of the first generated image. Holes to be inpainted cover up to 25% of the image. The authors note that reprojection only approximately places highlights, but they hypothesize this is sufficient for estimating intrinsic appearance maps because roughness and metallic depend more on the sharpness and contrast of highlights than on their exact position.

Each generated image is then passed to an SVBRDF predictor to obtain basecolor, roughness, and metallic maps. Finally, all per-view maps are merged into a scene-space texture atlas using the algorithm implemented in MeshLab, applied independently to each quantity, which blends each texel's observations according to geometric and color criteria rather than populating the atlas incrementally.

For the evaluation, the authors use pre-trained models from Zhu et al., Kocsis et al., and Zeng et al. (all trained on InteriorVerse) and train the UNet variants and GenPercept on InteriorVerse themselves. To train with hyperfeatures, they perform inverse DDIM sampling to invert each training image through Stable Diffusion and obtain the corresponding UNet activation features; hyperfeatures are extracted using 11 timesteps and a projection dimension of 384. The dataset is curated by denoising Monte-Carlo noise with Mitsuba's Optix integration, cropping to squares and resizing to 512², and removing files below 2 MB (4% of the dataset). Training uses a combination of an L1 loss in the FLIP perceptual color space and VGG LPIPS for basecolor, and plain L1 for metallic and roughness, with weights α_b = 1.0, α_m = 2.0, α_r = 0.5, and λ_b = λ_r = 0.5, using an Adam optimizer at a learning rate of 10e-5 for a maximum budget of 8 days on 4 GPUs (RTX 6000s or RTX 8000s). Evaluation uses scale-invariant metrics for basecolor and reports PSNR, SSIM, and LPIPS; consistency is measured with the Stop-the-Pop flickering metric, adapted to use pixel-perfect depth maps and cameras for warping, over 5 synthetic scenes not part of InteriorVerse with 100 views each.

Why This Matters

Impact on research: The work reframes generative material design as a pipeline that mixes image synthesis with prediction, and shows that a plain UNet can beat more elaborate recent designs on accuracy while a hyperfeature-augmented UNet wins on multiview consistency. It also provides a concrete, quantitative comparison of the accuracy-versus-coherence-versus-runtime tradeoffs that matter for texture atlases, which the paper argues is the only public dataset with the paired BRDF maps needed for this evaluation.

Real-world applications:

  • Rapid prototyping of relightable materials for game and film assets, where a few minutes per scene is acceptable but per-view minutes are not.
  • Interior design and virtual staging, where a designer iterates on text or image prompts over existing room geometry.
  • E-commerce and product visualization, rendering furniture and fixtures under arbitrary lighting from a merged atlas.
  • AR/VR and digital twins, where untextured scanned spaces need plausible, relightable appearance without manual material authoring.

Industry relevance: the pipeline relies on widely available components (Stable Diffusion v1.5, its ControlNets revision v1.1, the Diffusers Multi-ControlNet pipeline, MeshLab), runs in image space, and is fast enough for iteration, making it practical for content creation tools; two of the authors are affiliated with Adobe Research and MIT, and the work was funded by the ERC Advanced Grant NERPHYS with support from Adobe and NVIDIA.

Future Directions

  • Building a richer training dataset: the authors state InteriorVerse has a small number of materials for floor, walls, and furniture, metallic objects are rare, and 99% of roughness values are below 0.8, and expect a richer dataset to improve the variety of generated materials.
  • Extending the work to multiview SVBRDF extraction, which has been explored for photographs but may be challenged by the illumination inconsistencies of generated images.
  • Fixing reprojection artifacts for thin objects and small geometry, illustrated by a wooden fork whose material and geometry are not cleanly generated.
  • Reducing visible seams from image inpainting that can persist in the merged texture atlas, and improving prompt adherence using recent diffusion models to reduce the prompt engineering iterations currently required.

Target Audience

Researchers and practitioners in computer graphics, inverse rendering, and generative content creation who want an empirical comparison of SVBRDF predictors in a generative setting; technical artists and pipeline engineers building automated texturing tools for 3D scenes; and students interested in the interaction between diffusion-based image synthesis and physically based material estimation.

Authors’ abstract

Digital content creation is experiencing a profound change with the advent of deep generative models. For texturing, conditional image generators now allow the synthesis of realistic RGB images of a 3D scene that align with the geometry of that scene. For appearance modeling, SVBRDF prediction networks recover material parameters from RGB images. Combining these technologies allows us to quickly generate SVBRDF maps for multiple views of a 3D scene, which can be merged to form a SVBRDF texture atlas of that scene. In this paper, we analyze the challenges and opportunities for SVBRDF prediction in the context of such a fast appearance modeling pipeline. On the one hand, single-view SVBRDF predictions might suffer from multiview incoherence and yield inconsistent texture atlases. On the other hand, generated RGB images, and the different modalities on which they are conditioned, can provide additional information for SVBRDF estimation compared to photographs. We compare neural architectures and conditions to identify designs that achieve high accuracy and coherence. We find that, surprisingly, a standard UNet is competitive with more complex designs. Project page: http://repo-sam.inria.fr/nerphys/svbrdf-evaluation

Read the original paper