Skip to content
AI.info

Research

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Overview Research area: Computer vision / 3D generative modeling — specifically high-resolution image-to-3D generation and learned 3D shape representations (latent tokenization, rectified-flow generat

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
arXiv
2610.02201
Published
2026-10-01
Authors
Tianjiao Yu, Xinzhuo Li, Yifan Shen, Ying Shen, Kiet A. Nguyen, Adheesh Sunil Juvekar, Ismini Lourentzou

AI summary

Overview

Research area: Computer vision / 3D generative modeling — specifically high-resolution image-to-3D generation and learned 3D shape representations (latent tokenization, rectified-flow generative models, and topology-aware learning).

Technical level: Advanced. The paper assumes familiarity with latent 3D representations (voxels, triplanes, sparse voxel tokenizers), rectified flow / diffusion transformers, VAEs, and computational topology (persistent homology, Betti numbers, Morse theory).

Scope in one sentence: The paper proposes SILSA, a framework that represents 3D shapes with a fixed set of 384 overlapping "sliding-window slice latents" across three canonical axes, trained with topology-preserving losses and generated in a single rectified-flow stage using a shared volumetric anchor lattice.

What This Paper Is About

High-resolution 3D generation typically works by predicting which voxels or cells are occupied and then synthesizing local geometry inside them. That approach is effective but expensive, because it splits continuous surfaces into many small tokens whose count grows with surface area and part complexity, and it often produces shapes that look right on average while breaking thin supports, filling in holes, or merging separate components. The paper's goal is a representation that is both compact and structurally faithful: instead of many local voxel tokens, each token summarizes a whole planar cross-section (a "slice") of the shape, so a fixed, object-independent number of tokens can carry the shape's structure.

Key Contributions

  1. SILSA, a topology-aware image-to-3D framework that encodes shapes into a compact set of spatially grounded sliding-window slice latents along the x, y, and z axes. The representation uses a fixed 3N = 384 slice tokens (N = 128 bins per axis, window size w = 8), independent of object occupancy, surface area, or part complexity.

  2. A topology-aware SliceVAE that combines overlapping slice aggregation, sparse volumetric decoding, persistent-homology matching within slices, and Betti-transition supervision across neighboring slices, to preserve surface geometry, connected components, and hole structures.

  3. A single-stage rectified-flow generator with a Volumetric Anchor Lattice (VAL) — a persistent 3D feature grid that slice tokens read from and write to at each transformer block, enabling cross-axis coordination in a shared spatial workspace. Because the slice layout is fixed, generation requires no separate active-voxel prediction stage.

  4. A slice-wise topology-preserving loss consisting of a per-slice persistence-diagram matching term and an inter-slice Betti-transition matching term, avoiding the cost of full-volume topology matching.

Main Findings

  • Image-to-3D quality: SILSA achieves the best FD (10.16), PSNR (32.74), coverage (79.08%) and MMD (14.02‰) in Table 1, while matching the best KD (0.08) and LPIPS (0.05). PSNR improves from 30.12 to 32.74 over SparseFlex, an 8.7% relative gain, and coverage rises from 73.12% to 79.08%, a 5.96-point absolute improvement.

  • CLIP trade-off: SparseFlex attains a slightly higher CLIP score (88.22) than SILSA (87.94); the paper argues SILSA's substantially better geometric and distributional metrics indicate improved 3D fidelity without sacrificing image alignment.

  • VAE reconstruction quality: Compared with SparseFlex (the strongest baseline), SILSA reduces Chamfer Distance from 0.61 to 0.59, improves F-Score@0.01 from 96.18 to 96.79, improves F-Score@0.005 from 83.62 to 84.03, and increases IoU from 92.54 to 93.01.

  • Topology preservation: Betti-Err drops from 1.743 (SparseFlex) to 1.582 (SILSA), a 9.2% relative reduction, indicating better preservation of connected components and holes rather than only surface-level accuracy.

  • Token efficiency: SILSA uses 384 tokens, which the paper reports as 70.0% fewer than the next-most compact baseline (Dora, 1,280 tokens) and over 98% fewer than sparse or hierarchical tokenizers (Trellis: 19,847 ± 4,312; XCube: 64,821; SparseFlex: 87,453 ± 18,264).

  • Compute savings: Relative to Dora, training memory falls from 14.6 GB to 8.7 GB (a 40.4% reduction), training time from 0.38 s/iter to 0.21 s/iter, and inference from 0.82 s/shape to 0.34 s/shape (a 58.5% reduction). SILSA's model has 96M parameters and uses 1 stage, versus 2 stages for XCube, Trellis, and SparseFlex.

  • Cross-axis communication matters: With no communication between the three slice streams, CD rises to 5.87, IoU falls to 49.26, and Betti-Err rises to 17.91. Direct cross-attention gives CD 0.60, IoU 93.18, Betti 1.64; VAL gives the best result with CD 0.59, IoU 93.01, Betti 1.58.

  • Anchor lattice resolution: D = 16 (the default) beats D = 8 (CD 0.65, IoU 91.37, Betti 1.86) and D = 32, which slightly improves IoU to 93.18 but worsens CD to 0.60 and Betti-Err to 1.62.

  • Topology loss ablation: Removing L_topo gives CD 0.78, IoU 81.76, Betti 4.43; using only L_PH gives 0.72 / 84.41 / 2.91; only L_trans gives 0.69 / 87.16 / 2.87; the full L_topo gives 0.59 / 93.01 / 1.58.

  • Qualitative behavior: The paper reports that SILSA better preserves thin supports, articulated parts, dense branches, holes, and repeated structures, where competing tokenizers smooth fine details, merge nearby components, or distort fragile parts.

Methodology in Plain English

The core idea is to stop representing a 3D shape as a cloud of small voxel tokens and instead describe it with a small number of thick, overlapping planar cross-sections.

  • Encoding. A mesh is sampled into a point cloud with normals and divided into N = 128 bins along each of the x, y, and z axes. For each bin, the encoder gathers points within a sliding window of neighboring bins (w = 8), adds each point's offset from the bin center, passes them through a shared MLP, and max-pools the result. That pooled feature becomes the parameters of a Gaussian latent for that slice. The full shape is therefore 3N = 384 latents. Empty windows get a learned "empty" embedding.

  • Decoding. The three axis-wise latent sequences are scattered into a shared coarse 3D feature grid (slice latents that map to the same grid plane are averaged), refined by a sparse transformer decoder, and then upsampled through two self-pruning stages from 16³ to 256³. A final linear head predicts per-cell isosurface parameters (SDF values, vertex deformations, interpolation weights), and a mesh is extracted with differentiable Dual Marching Cubes.

  • Topology supervision. During VAE training, the decoder's SDF grid is converted to soft occupancy maps at Ns evenly spaced cross-sections per axis. Each cross-section is treated as a superlevel-set filtration whose persistence diagram records connected components (dimension 0) and holes (dimension 1). One loss matches predicted persistence diagrams to ground-truth ones (with unmatched features penalized against the diagonal); a second loss matches the sequence of Betti number changes between adjacent slices. Because Betti counts are discrete, they are computed on thresholded occupancy in the forward pass with a straight-through estimator for gradients.

  • Generation. After the VAE is trained and frozen, a rectified-flow transformer is trained to map noise to the ground-truth slice latents, conditioned on DINOv2 image features. The Volumetric Anchor Lattice is a persistent D×D×D×C grid: each slice token knows which plane it belongs to (x-slice at depth k maps to G[k',:,:], and so on, with k' = floor(kD/N)). Each transformer block does intra-axis self-attention, anchor read (cross-attention to the token's depth plane), gated anchor write, image cross-attention, and a feed-forward network. The lattice is reinitialized to zeros at each velocity evaluation and accumulates cross-axis evidence across the L stacked blocks.

  • Training and evaluation data. Training uses Trellis-500K. Evaluation uses 200 randomly sampled Toys4K assets and 50 in-the-wild images, with no overlap with the training set. Baselines for reconstruction are Dora, XCube, Trellis, and SparseFlex; image-to-3D baselines include Shap-E, LN3Diff, Direct3D, 3DTopia-XL, InstantMesh, GaussianAnything, XCube, Dora, SAR3D, Trellis, and SparseFlex.

Why This Matters

Impact on research. The paper reframes high-resolution 3D generation as a representation problem rather than purely a capacity problem. It argues that fixed-length, spatially grounded cross-sectional latents can replace variable-length active-voxel tokenization while improving both fidelity and topology, and it shows that topology supervision can be applied cheaply at the slice level instead of over full volumes. It also removes a whole pipeline stage: because the slice layout is fixed, no separate active-structure prediction is needed, which is a departure from the multi-stage design of prior voxel-latent systems.

Real-world applications (derived from the capabilities the paper demonstrates and motivates; the paper does not itself enumerate deployment scenarios):

  • Generating 3D assets from a single photograph for games, AR/VR, and e-commerce, where thin supports, railings, and repeated parts are common and commonly broken by current methods.
  • Reverse engineering or digitizing real objects where structural correctness — holes, handles, open topology — matters more than smooth surfaces.
  • CAD-adjacent and mechanical/articulated part modeling, where merging two nearby components or filling a hole is a functional failure, not a cosmetic one.
  • Scientific and botanical or biological shape reconstruction, where dense branching and connectivity (the paper's examples include plants with many branching stems) are the signal of interest.

Industry relevance. Efficiency claims are directly relevant to serving costs: 384 tokens versus tens of thousands, a 40.4% reduction in training memory (14.6 GB to 8.7 GB), a 58.5% reduction in inference time (0.82 s/shape to 0.34 s/shape), and a 96M-parameter model that is smaller than Dora (124M), Trellis (347M), and SparseFlex (213M). For any team running 3D generation at scale, that combination of a single stage, fixed sequence length, and lower memory makes throughput predictable in a way variable-length tokenizers are not.

Future Directions

  • Scaling resolution and slice count. The paper fixes N = 128 bins per axis and D = 16 anchors, and reports that D = 32 does not help. Whether the representation's advantage holds at substantially higher output resolution (beyond the 256³ decoding grid) is not reported.

  • Text-conditioned generation. All reported generation experiments are image-conditioned with DINOv2 features and 50 in-the-wild images; the paper does not report text-conditioned results, so whether slice latents generalize to prompt-based generation is an open question.

  • Better handling of the empty-window case. Empty windows receive a learned "empty" embedding, and the interaction between this and the fixed token budget for shapes with large empty regions is not analyzed.

  • Reducing dependence on topology losses at training time. The ablation shows Betti-Err jumps from 1.58 to 4.43 without L_topo; whether topology can be enforced at inference or via architectural priors rather than paired persistence-diagram supervision is unexplored.

  • Broader topology evaluation. Betti-Err is reported as an average mismatch in connected components and holes; how SILSA performs on higher-dimensional homology or on out-of-distribution shape categories is not reported.

Target Audience

Researchers and graduate students working on 3D generative models, neural 3D representations, and latent tokenization will get the most from this paper, particularly those already familiar with voxel-latent, sparse-voxel, and diffusion/flow-based 3D pipelines. It is also relevant to practitioners applying computational topology (persistent homology, Betti numbers) as training supervision, and to engineers building production image-to-3D systems who care about the trade-off between token count, memory, inference latency, and structural fidelity. Beginners will find the section-3 formalism (filtration notation, persistence diagrams, gated anchor writes) difficult without prior background, though the introduction and the qualitative figures are accessible on their own.

Authors’ abstract

High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by $8.7\%$, coverage by $5.96$ absolute points, and Betti error by $9.2\%$ over the strongest baseline, while using $70.0\%$ fewer tokens than the next-most compact baseline and over $98\%$ fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by $40.4\%$ and inference time by $58.5\%$. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.

Read the original paper