Research
LiTo: Surface Light Field Tokenization
Overview Research area: Computer vision and 3D generative modeling, specifically learned latent 3D representations that jointly encode object geometry and view-dependent appearance (surface light fiel

- arXiv
- 2603.11047
- Published
- 2026-03-11
- Authors
- Jen-Hao Rick Chang, Xiaoming Zhao, Dorian Chan, Oncel Tuzel
AI summary
Overview
- Research area: Computer vision and 3D generative modeling, specifically learned latent 3D representations that jointly encode object geometry and view-dependent appearance (surface light fields), plus single-image-to-3D generation.
- Technical level: Advanced. The paper assumes familiarity with Perceiver IO, flow matching, Gaussian splatting, spherical harmonics, Chamfer distance, FID/KID/PSNR/SSIM/LPIPS, and DINOv2 features.
- Scope (one sentence): The paper introduces LiTo, a tokenizer that compresses samples of a surface light field into a compact set of 3D latent vectors, and trains a latent flow-matching model on top of it to generate view-dependent 3D objects from a single image.
What This Paper Is About
Most existing learned 3D latent representations capture only geometry, or they capture appearance as view-independent diffuse color, so they cannot reproduce effects like specular highlights and Fresnel reflections that change with viewing angle. The authors build a latent representation from samples of the surface light field — a 5D function ℓ(x, d̂) mapping a surface point and a viewing direction to an outgoing RGB color — so that geometry and view-dependent radiance are encoded together in one latent space. They then learn the distribution of these latents conditioned on a single input image, so the generated 3D object's shape matches the input view and its appearance reflects the lighting and materials in the input.
Key Contributions
- A 3D latent representation that captures both geometry and view-dependent appearance by encoding surface light field information (surface point, viewing direction, color) into a compact set of 8192 tokens of dimension 32.
- A training framework that jointly supervises geometry and appearance using random subsamples of surface light field data drawn from RGB-depth multiview images, with geometry supervised through a flow-matching density model and radiance supervised by rendering Gaussian splats with spherical harmonics up to degree 3.
- A latent flow-matching generative model (a Diffusion Transformer with 623 million parameters) that learns the distribution of these latents conditioned on a single input image, producing full 3D objects whose appearance reflects the input lighting and materials.
- An encoder design that supports roughly 1 million input tokens using a 3D "patchification" approximation (K-nearest-neighbor assignment to queries) plus voxel-based self-attention with a half-cell shift per layer, making the large-sample input computationally tractable.
Main Findings
- Reconstruction quality (appearance): On Toys4k under TRELLIS's training lighting condition, LiTo reaches PSNR 34.16 ± 3.39, SSIM 0.985 ± 0.016, LPIPS 0.023 ± 0.018 at camera radius [3, 4], versus TRELLIS at 31.12 ± 3.39, 0.974 ± 0.022, 0.034 ± 0.022. At the harder camera radius [1, 3], LiTo reaches 32.36 ± 3.77, 0.967 ± 0.040, 0.055 ± 0.046 versus TRELLIS 27.57 ± 3.38, 0.941 ± 0.050, 0.090 ± 0.055. The paper states LiTo outperforms competitor appearance representations across all tested metrics.
- Geometry quality: In Chamfer distance (×10⁴, 100k points), LiTo with the mesh decoder obtains 87.17 ± 24.29 on PBR-Objaverse, 80.55 ± 27.59 on Toys4k, and 95.19 ± 23.64 on GSO. The paper reports that LiTo gives the best geometry among approaches that do not use a coarse geometry oracle and is competitive with those that do, while using a latent space the paper describes as 10× smaller.
- Geometry without an oracle: LiTo does not require ground-truth coarse occupancy for decoding, unlike TRELLIS and TripoSF, yet it outperforms most geometry-only latent representations (Shape Tokens: 126.0 ± 23.20, 119.8 ± 28.02, 130.5 ± 20.72 on the three datasets).
- Single-image-to-3D generation: On Toys4k with fixed area lighting matching TRELLIS and CFG scale 3.0 for both models, LiTo improves conditioning-view FID from TRELLIS's 12.84 to 6.219 and KID from 0.088 to 0.009, with FID_dino dropping from 84.692 to 41.621 and KID_dino from 2.311 to 1.333. CLIP score is 0.905 ± 0.041 for LiTo versus 0.899 ± 0.045 for TRELLIS.
- Novel-view generation does not degrade: Rendering four novel views at 30° pitch, LiTo reports FID 6.216, KID 0.058, FID_dino 66.530, KID_dino 3.522, compared with TRELLIS at 7.600, 0.100, 67.458, 3.166. The paper concludes that increased faithfulness to the input view does not significantly degrade overall generation performance.
- Input-view alignment: Because the model takes points rather than axis-aligned voxel grids as input, coordinate transforms can be applied during training. For each sample the world coordinate system is rotated so the input view's camera pose is identity orientation, so generated objects are consistently oriented relative to the input view — something the paper states TRELLIS does not do without post-processing.
- Spherical harmonic degree matters: Ablations show that increasing spherical harmonic degree from 0 to 3 consistently improves capacity in all three result tables (rows 1-3 to 1-6, 2-3 to 2-6, 3-3 to 3-6), whereas adding ray information alone (rows 1-2 vs. 1-3, etc.) does not directly improve appearance modeling; the paper hypothesizes that zero-degree harmonics cannot capture view-dependent effects and become a bottleneck.
- Camera encoding ablation: The authors report that although they considered explicit camera geometry encodings such as Plücker ray embeddings, in practice this reduced overall performance (Tab. S6).
Methodology in Plain English
The authors start from the observation that RGB-depth images taken from many viewpoints are effectively samples of an object's surface light field: back-projecting a depth map gives a surface point, the camera model gives a viewing direction, and the pixel gives a color. They box-normalize scenes to [-1, 1] and render 150 images at resolution 1036 × 1036 with a 40-degree field of view uniformly on a sphere of radius 3.5, yielding 160 million light field samples. Of these, they feed a random subset of 2²⁰ (about 1 million) samples into a Perceiver IO encoder and keep the remainder as ground truth for supervision.
To keep cross-attention affordable with a million inputs, they approximate the patchification used in Vision Transformers directly on 3D surfaces: they pick k = 8192 random samples as queries, assign every input sample to its nearest query by ℓ₂ distance, and let each query attend only to its assigned samples. For self-attention, tokens inside the same cell of a coarse voxel grid attend to each other, with the grid shifted by half a cell each layer. The encoder outputs 8192 latent tokens of dimension 32.
Supervision is indirect because only sparse samples of the field are available. Geometry is trained as a 3D probability distribution approximating a delta on the object's surface, using the flow-matching loss from Shape Tokens, which also lets the model sample surface point clouds and estimate normals at inference. Appearance is trained by decoding the latent into 3D Gaussians with degree-3 spherical harmonics and rendering images from random viewpoints, comparing against ground truth with an L2 loss plus an LPIPS term weighted by λ = 0.2. Training uses batch size 256 for 90k iterations on 64 GPUs for 9 days. Component sizes are 59.2 million parameters for the encoder, 8.8 million for the flow-matching velocity decoder, and 77.3 million for the Gaussian decoder, which emits 64 Gaussians per occupied voxel.
For generation, a 623-million-parameter Diffusion Transformer with a zero-initialized learnable positional encoding per token is conditioned on DINOv2-large image embeddings passed through a patchification layer. It is trained for 600k iterations with effective batch size 256 on 128 H100 GPUs for 20 days. Training data is the 500k high-quality object subset of Objaverse-XL selected by TRELLIS, split 8:1:1 into train/validation/test, with each object paired with three lighting conditions (fixed smooth area lighting matching TRELLIS, an all-white environment map, and randomly placed lights). Evaluation uses Toys4k, GSO, and a 200-object PBR-material subset of Objaverse-XL the authors call PBR-Objaverse.
Why This Matters
Impact on research. The paper argues that view-dependent appearance is largely missing from learned 3D latents: geometry-only methods capture shape but no radiance, and appearance-aware methods such as TRELLIS mean-pool multiview features, discarding angular variation. LiTo shows that feeding viewing direction into the encoder and decoding higher-order spherical harmonics produces better visual quality "without significant degradation in geometric accuracy," and that a latent encoding complete object information can support single-stage generation, unlike two-stage pipelines that require coarse occupancy to be known in advance.
Real-world applications.
- Single-image-to-3D asset creation for games, film, and simulation pipelines, where a generated asset inherits the lighting and material look of the reference photo.
- Product visualization and e-commerce, where reflective and metallic objects must look correct from arbitrary viewpoints rather than as flat diffuse textures.
- Augmented and virtual reality content, where the object must align with the user's input viewpoint immediately after generation.
- Physically based rendering asset authoring, where roughness/metallicity-style outputs (as in 3DTopia-XL's PrimX) matter for downstream renderers — the paper positions view-dependent modeling as the missing piece for realistic appearance.
Industry relevance. The work originates from Apple and the acknowledgements credit Apple infrastructure; the broader trend it participates in — compact 3D latents that a transformer can generate in one stage — is directly relevant to any company producing 3D content at scale. The paper notes that its Gaussian decoder predicts degree-3 spherical harmonics rather than view-independent color, and that its training framework avoids the watertight-mesh, mesh-to-field, or radiance-field optimization steps that many geometry latents require. The paper does not report commercial deployments or product integrations.
Future Directions
- Transparency and high-frequency specularity: The stated limitation is that the 3DGS implementation supports spherical harmonics up to degree 3, which limits faithful reconstruction of transparent materials or high-frequency specularities. A higher-degree or alternative radiance representation is a natural next step.
- Better use of viewing direction: Ablations show that adding ray input alone did not improve appearance modeling, which the authors attribute to zero-degree harmonics acting as a bottleneck. Understanding how to make directional input pay off more fully remains open.
- Close-up fidelity: The supplementary results note that close-up views demand greater high-frequency detail and that all methods face challenges there, though LiTo is reported as most robust. Improving close-range reconstruction is an open direction.
- Camera encoding and generation conditioning: Since Plücker ray embeddings reduced performance in their ablation, the question of how best to inject camera geometry into the generative model is unresolved.
- Orientation and controllability at generation time: The model achieves input-view alignment by rotating the world so the input camera is identity, which makes orientation inference unnecessary. Alternatives that retain explicit orientation control are not explored.
Target Audience
Researchers and practitioners in 3D generative modeling, neural rendering, and inverse graphics who are already comfortable with latent set representations, flow matching, and Gaussian splatting. It is also relevant to graphics and content-creation engineers who need generated 3D assets with correct view-dependent material appearance, and to readers comparing latent 3D representations (TRELLIS, 3DTopia-XL, TripoSG, TripoSF, Shape Tokens, CLAY, Hunyuan3D) on appearance, geometry, latent size, and preprocessing requirements. Readers without a computer vision or graphics background will find the notation and the metric-heavy results sections difficult.
Authors’ abstract
We propose a 3D latent representation that jointly models object geometry and view-dependent appearance. Most prior works focus on either reconstructing 3D geometry or predicting view-independent diffuse appearance, and thus struggle to capture realistic view-dependent effects. Our approach leverages that RGB-depth images provide samples of a surface light field. By encoding random subsamples of this surface light field into a compact set of latent vectors, our model learns to represent both geometry and appearance within a unified 3D latent space. This representation reproduces view-dependent effects such as specular highlights and Fresnel reflections under complex lighting. We further train a latent flow matching model on this representation to learn its distribution conditioned on a single input image, enabling the generation of 3D objects with appearances consistent with the lighting and materials in the input. Experiments show that our approach achieves higher visual quality and better input fidelity than existing methods.