Skip to content
AI.info

Research

Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?

Overview Research area: Computer vision / 3D generative modeling, specifically native 3D texture generation (predicting surface colors directly in 3D space for a given geometry). Technical level: Inte

Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
arXiv
2609.34621
Published
2026-09-28
Authors
Jiangshan Wang, Zeqiang Lai, Jiayi Guo, Xin Yang, Xin Huang, Jiarui Chen, Ziheng Ouyang, Chunchao Guo, Xiangyu Yue

AI summary

Overview

  • Research area: Computer vision / 3D generative modeling, specifically native 3D texture generation (predicting surface colors directly in 3D space for a given geometry).
  • Technical level: Intermediate. The core idea is intuitive, but the paper assumes familiarity with VAEs, diffusion/flow-matching transformers (DiTs), sparse 3D convolutions, and latent-space conditioning.
  • Scope: The paper asks whether a high-fidelity native 3D texture generation framework can be trained entirely on constructed 2D image data instead of real textured 3D assets, and presents a framework called Tex-Zero that does so.

What This Paper Is About

Native 3D texture generation models are believed to require large-scale, high-quality real 3D assets, which are scarce and expensive to acquire. The authors observe that only the fine-grained color information in training data is essential, while the geometry can be manually constructed rather than captured from real objects. They therefore convert ordinary 2D images into synthetic 3D training samples — each image becomes a plane in 3D space whose patches are randomly rotated and aggregated — and train both a VAE and a DiT exclusively on that data, then test them on real 3D assets.

Key Contributions

  1. A data construction pipeline that transforms large-scale 2D images into effective training data for native 3D texture generation, showing that the geometric structures of training data need not be semantically meaningful.
  2. A unified 2D–3D VAE (Tex-Zero VAE) trained only on image-derived data that reconstructs both real 3D textures and 2D images, providing a shared latent space for the two modalities.
  3. A DiT design (Tex-Zero DiT) that represents conditioning multi-view images as planes in 3D space and encodes them with the same Tex-Zero VAE, eliminating the representation gap between conditions and target textures.
  4. A demonstration that Tex-Zero generates high-fidelity 3D textures with fine-grained details using only images, outperforming representative baselines trained on real textured 3D data.

Main Findings

  • Training without 3D assets works: The Tex-Zero VAE never observes real 3D assets during training, yet reconstructs real 3D assets with competitive quality — e.g., at f8c16 it reaches LPIPS 0.0154, PSNR-PC 34.03, PSNR 42.41, and SSIM 0.987 on 3D asset reconstruction.
  • Best 2D image reconstruction among compared VAEs: At f8c16, Tex-Zero achieves LPIPS 0.0217, PSNR 39.82, and SSIM 0.956 on 2D images, versus 0.1430 / 36.71 / 0.944 for the 3D-asset-trained Sparse VAE at the same f8c16 configuration and 0.0209 / 39.63 / 0.933 for FLUX VAE (image-trained, with no 3D asset reconstruction reported).
  • 3D VAEs trained on 3D data struggle with images: NaTex VAE reaches LPIPS 0.1977 / PSNR 34.36 / SSIM 0.918 and TRELLIS.2 VAE reaches 0.1487 / 33.81 / 0.915 on 2D image reconstruction, far behind Tex-Zero.
  • Generation beats baselines: In six-view evaluation, Tex-Zero scores LPIPS 0.0340, PSNR 35.68, SSIM 0.983 versus NaTex at 0.0754 / 27.74 / 0.949. In front-view evaluation, Tex-Zero scores 0.0294 / 35.31 / 0.985 versus NaTex at 0.0669 / 27.49 / 0.947 and TRELLIS.2 at 0.1187 / 21.38 / 0.900. TRELLIS.2 six-view results are not reported.
  • Image data still helps when 3D data exists: Joint 2D+3D training improves both the VAE (LPIPS 0.0140, PSNR 42.85, SSIM 0.988) and the DiT (LPIPS 0.0216, PSNR 37.68, SSIM 0.983) compared to 3D-only training (VAE 0.0343 / 41.84 / 0.981; DiT 0.0302 / 36.23 / 0.980) at the f16c32 configuration.
  • Every preprocessing step matters (20K-step ablation): A fixed plane with no rotation and no aggregation yields LPIPS 0.2160, PSNR-PC 11.34, PSNR 18.16, SSIM 0.777. Adding global rotation gives 0.1107 / 20.22 / 28.38 / 0.904; 4 independently rotated patches give 0.0591 / 26.83 / 35.05 / 0.963; adding aggregation gives 0.0513 / 27.81 / 36.18 / 0.966; 16 aggregated patches give the best result at 0.0355 / 28.97 / 37.16 / 0.973.
  • Unified latent space is the key conditioning design: With 20K training steps, DINO as image encoder gives LPIPS 0.1228 / PSNR 25.67 / SSIM 0.886; a separate image VAE gives 0.0838 / 29.65 / 0.955; a shared VAE trained on 3D assets (unified latent) gives 0.0835 / 30.42 / 0.966; the Tex-Zero VAE (unified latent, image-trained) gives the best 0.0481 / 32.54 / 0.970.
  • Warm-up is required: Without a warm-up stage in which each image is a single randomly rotated colored plane, the VAE fails to converge.
  • Theoretical justification: The appendix argues that fixed-plane training leaves 3D convolution kernel directions unsupervised (gradient zero for offsets with Δz ≠ 0), globally rotated planes remain coplanar (minimum eigenvalue of the coordinate covariance is zero), and only spatial aggregation brings differently oriented patches within the VAE's receptive radius so their features interact.

Methodology in Plain English

The authors start from the observation that a 2D image and a 3D texture share the same basic structure: colors assigned to spatial locations. Their pipeline turns each image into a synthetic 3D sample:

  1. Image as a colored plane. Every pixel becomes a 3D point. Pixel coordinates set the position, RGB values set the color, all points lie on the plane z = 0, and every point shares the normal (0, 0, 1).
  2. Patch-wise rotation. The image plane is cut into non-overlapping patches, and each patch is independently rotated by a random 3D rotation matrix and translated to a new location. Points and normals rotate together; colors stay unchanged, so the picture content inside each patch is preserved.
  3. Aggregation. Rotated patches are pulled into a bounded 3D region so they overlap or intersect rather than float apart, creating occlusion patterns and complex local geometry.
  4. Voxelization and rendering. Transformed points are assigned to a voxel grid (one point kept per voxel), and the colored point cloud is rendered from canonical orthographic viewpoints to produce conditioning images and foreground masks. Each image thus yields a tuple of geometry, colors, and multi-view conditions.

VAE training: The Tex-Zero VAE follows a FLUX VAE-like architecture with the encoder and decoder implemented as sparse 3D operations. It is trained with a pointwise color reconstruction loss, an image-space perceptual loss (LPIPS, computed after inverting the patch transformations and reassembling the original 2D layout), and KL regularization. A warm-up stage using a single rotated plane precedes the full patch-based pipeline.

DiT training: The Tex-Zero DiT takes a noisy texture latent, six canonical view images represented as colored planes and encoded by the frozen Tex-Zero VAE, plus geometry features from surface normals. It uses a joint 4D rotary positional embedding with a group index (0 for target texture tokens, 1–6 for front, right, back, left, top, bottom views) and a standard flow-matching objective. A random subset of views is sampled per step, and image-condition dropout enables classifier-free guidance. Training uses only constructed image data.

Data: Both models are trained on SA-1B, BLIP3o-60k, and ShareGPT-4o, roughly 11.1 million images total, resized to 1536 × 1536. The default DiT configuration uses the Tex-Zero VAE with a spatial downsampling factor of 16 and 16 latent channels. Comparisons include baselines trained on approximately one million in-house textured 3D assets. Evaluation uses PSNR, SSIM, LPIPS on rendered views and PSNR-PC on the point cloud.

Why This Matters

Impact on research: The paper challenges the assumption that native 3D texture generation requires real 3D assets, and suggests the geometric component of training data can be synthetic while appearance comes from abundant 2D images. If the finding generalizes, it removes a major bottleneck for scaling 3D generative models by letting them draw on the same kind of large-scale image data that has driven progress in 2D.

Potential real-world applications (the paper itself does not enumerate applications; these are plausible domains that would benefit from cheap, high-fidelity texturing):

  • Game development, where texturing large libraries of 3D assets is labor-intensive.
  • AR/VR content creation and digital-twin modeling.
  • E-commerce 3D product visualization, where product photos could be mapped onto product geometry.
  • Animation and visual effects pipelines that need consistent textures across many assets.

Industry relevance: Producing textured 3D data at scale is expensive and requires specialized equipment. A method that turns existing image corpora into usable supervision lowers the barrier to building 3D texture models, and the joint-training result shows the approach also adds value on top of existing in-house 3D datasets rather than replacing them.

Future Directions

  • Testing whether the constructed-data paradigm extends to other 3D generation tasks beyond texture, such as geometry or full asset generation.
  • Investigating scalability: the paper uses roughly 11.1 million images and an approximately one-million-asset in-house 3D set, but does not report how performance scales with image data volume.
  • Understanding precisely which properties of the constructed geometry drive transfer to real assets, beyond the ablations over patch count, rotation, and aggregation.
  • Reducing reliance on the warm-up stage and on careful data preprocessing design, since the VAE fails to converge without warm-up and the number of patches strongly affects results.

Target Audience

Researchers and engineers working on 3D generative models, texture synthesis, and latent-space representation learning, as well as practitioners building asset pipelines who need textured 3D content at scale. Readers with background in VAEs, diffusion/flow-matching models, and 3D representations will get the most out of the methodology and ablations.

Authors’ abstract

Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view reference images. It is generally believed that training such models requires large-scale, high-quality real 3D asset data, whose acquisition remains a long-standing and challenging problem. In this work, we propose Tex-Zero, demonstrating that a high-fidelity native 3D texture generation framework can be trained without 3D assets. Our key observation is that only high-quality and fine-grained color information is essential for 3D texture training, while the required geometric information is less critical and can be manually constructed rather than obtained from real 3D assets. This finding makes it possible to transform abundant, high-quality 2D images into effective training samples for 3D texture generation. Specifically, we convert high-quality 2D images into 3D training samples by representing each image as a plane in 3D space and applying patch-wise random rotations and aggregation to construct complex geometric structures. Using these constructed image data, we train the Tex-Zero VAE, which can reconstruct real 3D assets with high quality despite never observing them during training. Building upon the Tex-Zero VAE, we train the Tex-Zero DiT also exclusively on the constructed image data, where the conditioning 2D multi-view images are transformed into planes in 3D space and also encoded by the Tex-Zero VAE, thereby reducing the representation gap and improving generation quality. Extensive experiments show that Tex-Zero generates high-fidelity 3D textures with fine-grained details solely using images as training data, offering a promising perspective on the data paradigm for scaling 3D texture generation.

Read the original paper