Skip to content
AI.info

Research

GOATex: Geometry & Occlusion-Aware Texturing

Overview Research area: Computer vision, specifically text-guided 3D mesh texturing with diffusion models, with a focus on occluded interior geometry. Technical level: Advanced. The paper assumes fami

GOATex: Geometry & Occlusion-Aware Texturing
arXiv
2511.23051
Published
2025-11-28
Authors
Hyunjin Kim, Kunho Kim, Adam Lee, Wonkwang Lee

AI summary

Overview

  • Research area: Computer vision, specifically text-guided 3D mesh texturing with diffusion models, with a focus on occluded interior geometry.
  • Technical level: Advanced. The paper assumes familiarity with diffusion models, depth-conditioned ControlNet, multi-view rendering, UV maps, ray casting, and mesh topology.
  • Scope: The paper introduces GOATex, a ray-based occlusion-aware framework that textures both the exterior and the occluded interior surfaces of a 3D mesh in ordered visibility layers, without fine-tuning a pretrained diffusion model.

What This Paper Is About

Existing text-to-texture methods for 3D meshes work well on surfaces visible from outside, but they have no mechanism for reaching geometry that is hidden inside an object, which leaves interiors untextured, seamy, or filled in with heuristics. GOATex treats the mesh as a layered structure, uses multi-view ray casting to assign each region a "hit level" describing its relative depth, then progressively reveals and textures each layer from outermost to innermost. The goal is a single, seamless, high-fidelity UV texture covering both the visible shell and the previously unreachable interior.

Key Contributions

  1. The authors introduce and address the task of generating realistic textures for occluded interior surfaces of 3D meshes alongside exterior regions, which they describe as a practically important but underexplored challenge in the mesh texturing literature, and claim GOATex as the first occlusion-aware 3D mesh texturing method targeting it.
  2. They propose GOATex, a ray-based occlusion-aware framework that textures exterior and interior regions without tuning pretrained diffusion models, and that supports dual prompting for inner and outer surfaces to give fine-grained control over layered structure.
  3. They contribute a two-stage visibility control strategy (residual face clustering plus normal flipping and backface culling) that progressively exposes interior geometry while preserving the object's global shape.
  4. They contribute a soft, visibility-weighted UV-space blending scheme that merges per-layer textures using view-dependent visibility confidence, and report user studies and GPT-based evaluations showing strong preference for GOATex over existing methods.

Main Findings

  • Qualitative superiority on interiors: TEXTure and SyncMVD rely on view-based generation plus unprojection, so they lack access to occluded geometry and fall back on heuristic filling such as Voronoi-based extrapolation, producing simplistic, inconsistent textures and visible seams. Paint3D and TEXGen perform better inside because they operate in UV space, but they still produce low-frequency textures such as flat colors or repetitive patterns because they cannot differentiate interior from exterior surfaces in the UV map.
  • Dual prompting works: Because each hit level is textured independently through the text-guided MVD module, a separate prompt can be assigned to the outermost layer (hit level 1) to control exterior appearance and to deeper layers (hit levels >= 2) to govern interiors, producing stylistically distinct and contextually appropriate interior textures.
  • Human raters favor GOATex consistently: In A/B preference tests, GOATex achieves strong and consistent preference from human raters across all comparisons.
  • GPT raters show a smaller margin against project-and-inpaint baselines: GPT-based evaluations rate GOATex's advantage lower when compared with TEXTure and SyncMVD. The authors speculate GPTs favor smoother interiors, such as leaving the interior unpainted or bleeding exterior textures inward, over methods that explicitly paint interiors (Paint3D, TEXGen, and GOATex), suggesting GPTs may emphasize surface smoothness more than interior completeness and consistency.
  • Human-GPT agreement statistics: Pearson correlations between GPT-based and human ratings were 0.22 for GPT-4o-mini, 0.31 for GPT-4o, 0.43 for GPT-4.1, and 0.34 for GPT-o3. Cohen's kappa values averaged over kappa > 0 were 0.54 for GPT-to-GPT, 0.31 for user-to-user, and 0.27 for user-to-GPT. In 17 individual evaluation cases user-to-GPT agreement exceeded kappa = 0.5, and in two cases GPT-o3 achieved perfect agreement (kappa = 1.0) with a human rater.
  • Ablation results (win rate over SyncMVD baseline, GPT-based A/B, columns 4o-mini / 4o / 4.1 / o3 / average): Hit Level Assignment 82.50 / 66.67 / 75.00 / 77.50 / 75.68; Superface Construction 84.62 / 70.00 / 75.86 / 89.74 / 80.27; Soft UV Merging 79.49 / 82.05 / 90.00 / 95.00 / 86.49; Residual Face Clustering 77.50 / 72.50 / 88.89 / 84.84 / 81.17; Normal Flipping & Backface Culling (the full method) 86.84 / 92.31 / 86.67 / 97.50 / 91.16.
  • Residual clustering alone can hurt: A slight drop is observed when residual face clustering is applied without normal flipping and backface culling, because geometrically adjacent faces can be grouped into the same hit level even when one fully occludes the other from all external viewpoints, so the occluded face is never textured and is excluded once that level completes.
  • FID/KID/CLIP metrics were deliberately not used: The authors state these metrics are ill-suited here because ground-truth textures are often unavailable, especially in occluded regions, making reference-based metrics like FID/KID and CLIP-I inapplicable, and because CLIP-based models are sensitive to rendering artifacts and correlate poorly with human judgment in texture assessment. The exact preference percentages shown in Figure 6 are not reported as numbers in the provided text.

Methodology in Plain English

The pipeline has five stages, as shown in Figure 2.

  1. Superface construction. Instead of assigning depth labels to every individual triangle, which would fragment flat regions because meshes contain many small faces, the authors use the Xatlas library to oversegment the mesh into connected, low-curvature regions called atlases. Each atlas becomes a "superface."
  2. Hit level assignment. Rays are cast from multiple viewpoints, and the intersection order of each ray with each face is recorded. Each ray's influence is weighted by the cosine similarity between the ray direction and the face normal, so near-head-on intersections count more. The weighted votes are summed per superface, and the most dominant intersection order becomes that superface's hit level. This sorts surfaces from outermost (lowest hit level) to innermost (highest).
  3. Visibility control. Rendering only the faces uniquely assigned to a level produces sparse, fragmented depth maps, which are out-of-distribution inputs to a diffusion model. Two fixes are applied. Residual face clustering renders the full set of still-untextured faces remaining from earlier levels, like peeling an onion, rather than only the newly assigned ones. Normal flipping with backface culling keeps all faces but flips the normals of already-textured ones, so previously front-facing surfaces are culled away while previously hidden back-facing surfaces rotate into view, exposing interiors without distorting the mesh's global shape.
  4. Per-level texturing. For each hit level, multi-view depth maps are rendered from the visibility-controlled face cluster and passed to a pretrained depth-conditioned multi-view diffusion (MVD) module to synthesize a distinct UV texture map for that level. No fine-tuning or training data is required.
  5. Soft UV-space blending. The same mesh region can be textured at multiple levels with different visibility confidence, so naive overwriting or uniform averaging creates artifacts and destroys inner/outer boundaries. The authors compute a per-texel UV-space weight from the absolute cosine similarity between the view direction and the face normal, aggregate across views, normalize across hit levels with a modified softmax, and use those normalized weights to combine the per-level textures.

Implementation specifics: Stable Diffusion 1.5 with a depth-based ControlNet generates multi-view images from text prompts augmented with view-specific cues. Each view is rendered at 768 x 768 with a 96 x 96 latent resolution. The latent UV texture map is 512 x 512 and the final RGB UV texture map is 1024 x 1024. The maximum hit level is 4, and 16 hemispherical views are used (8 equatorial at 45 degrees, 8 elevated at 45 degrees). Rendering uses PyTorch3D. For hit level assignment, 17 cameras are used (the original texture-generation views plus an additional top-view camera), with ray casting at 1536 x 1536 resolution using the Open3D library. Experiments ran on a single RTX A6000 GPU requiring 12 GB of memory per inference; inference time is the number of hit levels multiplied by the inference time of the texture synthesis model.

Evaluation setup: The dataset consists of 139 assets from Objaverse and 87 from Objaverse-XL, giving 226 high-quality meshes with detailed interior geometry, each normalized to a unit bounding box. Objects come from 12 categories: box, bucket, bus, cabinet, car, drawer, house, lamp, room, shelf, tent, and truck. Captions were generated with GPT-4o: first, renderings at four hit levels are analyzed to identify key visual elements, then 10 diverse one-sentence prompts are generated per mesh describing material, texture, style, and pattern without naming colors or describing overall shape. Baselines are the project-and-inpaint methods TEXTure and SyncMVD and the UV-inpainting methods Paint3D and TEXGen; since TEXGen requires an initial UV map aligned with the geometry, the authors initialize it with the SyncMVD output UV map.

Why This Matters

Impact on research. The paper reframes mesh texturing as a visibility and layering problem rather than purely a 2D unprojection problem, and shows that a pretrained depth-conditioned diffusion model can texture occluded geometry with no fine-tuning. It also documents the mismatch between GPT-based aesthetic judgments and human raters on interiors, and reports correlation and agreement statistics that benchmark how well current GPT evaluators track human preference.

Real-world applications (as described in the paper):

  • Architectural visualization: house facades plus interior walls, doors, and furniture need detailed textures because users may inspect them closely during a VR walkthrough.
  • Vehicle simulation: the exterior body of a car or bus must look authentic, and interiors such as dashboards, seats, and ceiling panels need high fidelity for a seamless, immersive experience.
  • Gaming and animation asset creation, where high-quality mesh textures enhance believability of virtual environments.
  • Augmented reality and virtual reality experiences, where users move around, interact with, and change viewpoint on 3D assets and therefore see inner details.

Industry relevance. The method requires no task-specific fine-tuning or training dataset and runs within 12 GB per inference on a single RTX A6000, which keeps it cheap to adopt alongside existing diffusion pipelines. Separate prompts for inner and outer surfaces give artists and developers a new controllability axis, useful for stylized or layered asset design. The authors include KRAFTON AI and NC AI affiliations, both game-industry organizations.

Future Directions

  • Extending the framework to dynamic scenes and integrating material properties beyond texture.
  • Improving blending strategies for more complex topologies.
  • Adding semantic-aware refinement to hit-level assignment, for example grouping superfaces that belong to the same semantic volume using pretrained part-segmentation models, to improve cross-region consistency.
  • Integrating the soft blending directly into the denoising process so that multi-view and multi-hit-level texturing happen simultaneously for tighter cross-level coherence.

The paper also names a primary limitation: hit-level assignment is determined purely by geometric visibility (ray-intersection depth) and does not account for semantic coherence, so in complex geometries such as objects with thin openings or nested cavities, semantically unified regions may be split across multiple hit levels and cause minor texture discontinuities at their boundaries. In practice, the authors state that residual face clustering, view-dependent normal flipping, and soft UV-space blending mitigate most of these artifacts.

Target Audience

Researchers and graduate students working on 3D generative models, neural rendering, and diffusion-based asset creation will get the most from this paper, since it proposes a new task framing and a modular pipeline rather than a new network architecture. Practitioners in game development, VFX, simulation, and AR/VR who need textured assets with believable interiors also benefit, as the method requires no fine-tuning and fits an existing Stable Diffusion plus ControlNet workflow. Readers evaluating GPT-as-judge methodology will find the correlation and kappa analyses useful. Beginners would need background in diffusion models, depth-conditioned ControlNet, multi-view rendering, and mesh UV parameterization to follow the technical sections.

Authors’ abstract

We present GOATex, a diffusion-based method for 3D mesh texturing that generates high-quality textures for both exterior and interior surfaces. While existing methods perform well on visible regions, they inherently lack mechanisms to handle occluded interiors, resulting in incomplete textures and visible seams. To address this, we introduce an occlusion-aware texturing framework based on the concept of hit levels, which quantify the relative depth of mesh faces via multi-view ray casting. This allows us to partition mesh faces into ordered visibility layers, from outermost to innermost. We then apply a two-stage visibility control strategy that progressively reveals interior regions with structural coherence, followed by texturing each layer using a pretrained diffusion model. To seamlessly merge textures obtained across layers, we propose a soft UV-space blending technique that weighs each texture's contribution based on view-dependent visibility confidence. Empirical results demonstrate that GOATex consistently outperforms existing methods, producing seamless, high-fidelity textures across both visible and occluded surfaces. Unlike prior works, GOATex operates entirely without costly fine-tuning of a pretrained diffusion model and allows separate prompting for exterior and interior mesh regions, enabling fine-grained control over layered appearances. For more qualitative results, please visit our project page: https://goatex3d.github.io/.

Read the original paper