Skip to content
AI.info

Research

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering Overview Research area: Computer vision and multimodal machine learning, specifically 3D scene understanding and spatial reasoning w

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
arXiv
2609.38177
Published
2026-09-29
Authors
Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong

AI summary

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Overview

Research area: Computer vision and multimodal machine learning, specifically 3D scene understanding and spatial reasoning with Multimodal Large Language Models (MLLMs).

Technical level: Advanced. The paper builds on multi-view MLLM architectures, 3D Gaussian Splatting, feed-forward reconstruction models, and knowledge distillation, and assumes familiarity with transformer-based vision-language models.

Scope in one sentence: The paper proposes Imagine3D-LLM, a 7B multi-view MLLM that learns to assemble a compact 3D Gaussian Splatting reconstruction of a scene via a small set of learnable "Gaussian summary tokens" before answering questions about that scene, and reports gains across seven spatial reasoning and 3D understanding benchmarks.

What This Paper Is About

MLLMs handle single images well, but when they are given several views of the same scene and asked to reason about 3D structure, they fall well short of human performance. Prior attempts to close this gap feed the model fine-grained, pixel-level 3D signals, such as point-cloud coordinates attached to image patches, explicit visual markers, correspondence-supervised features, or features fused from 3D reconstruction foundation models such as CUT3R and VGGT. The authors argue, based on cognitive studies of human spatial perception, that humans instead identify recurring objects across views, use them as anchors to infer coarse relative geometry, and mentally assemble an approximate layout. Imagine3D-LLM is designed to do the analogous thing: build a compact, abstract 3D representation of the scene and condition its answer on that representation.

Key Contributions

  1. Gaussian summary tokens inside the LLM. A compact set of learnable tokens is inserted between the image tokens and the text tokens of a multi-view MLLM. Their hidden states at a middle LLM layer are decoded into 3D Gaussian Splatting primitives through a lightweight MLP head, so the model "imagines" a compact 3D scene before generating an answer.

  2. A bottleneck that induces object-level grouping. The number of Gaussian summary tokens is deliberately much smaller than the number of image tokens (M ≪ K·N′). This forces overlapping content observed across views to be merged into shared tokens, and the authors show via K-means clustering of the trained tokens that individual clusters localize single objects without any explicit clustering supervision.

  3. Joint reconstruction + language training, accelerated by distillation. The model is trained with a standard next-token prediction loss, a photometric reconstruction loss rendered at the input viewpoints, and a distillation loss from a frozen pretrained compact Gaussian teacher (ZipSplat) applied at both token and Gaussian-parameter level. The distillation is introduced because pure photometric supervision through the LLM converges too slowly to be practical.

  4. Evidence that reconstruction reshapes the LLM's own image features. Although only the summary tokens receive direct reconstruction supervision, the authors show sharper cross-view attention and cleaner, more semantically organized PCA visualizations of the LLM's image features after joint training, plus state-of-the-art results on seven benchmarks.

Main Findings

  • State-of-the-art across seven benchmarks. On SQA3D test, Imagine3D-LLM reaches 63.8 EM and 66.4 EM-R; on Real-3DQA test, 39.2 EM and 44.3 EM-R; on ScanQA val, 109.3 CIDEr, 20.5 BLEU-4, 21.3 METEOR, 49.7 ROUGE, and 29.9 EM. On Scan2Cap val (IoU@0.5) it reports 67.6 ROUGE, 48.3 BLEU-4, 32.6 METEOR, and 99.2 CIDEr; on ScanRefer val, 62.8 Acc@0.25 and 56.3 Acc@0.5; on Multi3DRefer val, 60.2 F1@0.25 and 55.0 F1@0.5. The strongest prior 3D LMM, Ross3D, scores 63.0/65.7 on SQA3D, 36.6/41.5 on Real-3DQA, 107.0 CIDEr, 17.9 BLEU-4, 20.9 METEOR, 50.7 ROUGE, 30.8 EM on ScanQA, 66.9/43.4/30.3/81.3 on Scan2Cap, 61.1/54.4 on ScanRefer, and 59.6/54.3 on Multi3DRefer.

  • SPAR-Bench results. Imagine3D-LLM-7B scores 68.5 average, with 60.5 low-level, 67.0 medium-level, and 76.0 high-level. The paper reports this surpasses the largest 72B-scale general-purpose model by over 29 points despite a 7B backbone, and outperforms the strongest spatially-aware baseline, 3DThinker-7B (63.3), by 5.2 points overall. Improvements are described as especially pronounced on the medium- and high-level splits.

  • Gains come from the objective, not extra data. Against a controlled baseline sharing identical backbone, training data, and schedule, Imagine3D-LLM improves on every benchmark: SQA3D 63.8 vs 56.5, ScanQA 29.9 vs 26.2, Scan2Cap 67.6 vs 63.1, ScanRefer 62.8 vs 58.3, Multi3DRefer 60.2 vs 57.4, SPAR 68.5 vs 60.9. The paper reports a +15.5 point improvement on the medium-level cross-view split of SPAR-Bench (from 51.5 to 67.0).

  • Middle-layer decoding is best. Extracting Gaussian tokens from layer ℓ = 14 of the 28-layer LLM gives 63.8 on SQA3D test and 68.5 on SPAR-Bench, versus 60.7/60.2 at layer 7 (early) and 61.7/64.0 at layer 21 (late).

  • Distillation accelerates convergence. Without distillation ("Recon. only"), the model scores 57.7/57.6 at 1 epoch, 61.9/65.1 at 2 epochs, and 63.7/67.9 at 4 epochs, nearly matching the full 1-epoch model at 63.8/68.5. The language-modeling-only baseline scores 56.5/60.9 at 1 epoch, 57.2/61.5 at 2 epochs, and drops to 52.3/55.2 at 4 epochs, which the authors attribute to overfitting.

  • Neither the teacher's tokens nor distillation alone help. Feeding the teacher's query tokens directly to the LLM gives 56.4/60.8 and supervising summary tokens with distillation but no reconstruction gives 56.6/60.5 — both on par with the 56.5/60.9 baseline. The reconstruction objective itself is what drives the improvement.

  • Token count matters, with a sweet spot. M = 1296 gives 61.9/64.6; M = 2592 (the chosen setting) gives 63.8/68.5; M = 5184 gives 60.3/58.4.

  • Emergent cross-view and object-level structure. Attention maps between a query point in one view and other views are sharper and more spatially focused than the baseline's, and PCA visualizations of image-token hidden states show corresponding objects encoded with consistent colors across views. The authors also report steady improvement in cross-view correspondence accuracy during training (detailed in the paper's Appendix B).

Methodology in Plain English

The team starts from LLaVA-Video-7B, a multi-view MLLM that turns each image into 210 tokens and feeds them to a language model together with the text instruction. They add a new block of learnable tokens — the Gaussian summary tokens — placed after the image tokens and before the text tokens, so each summary token can attend to all the visual evidence and choose what to encode. Crucially, there are far fewer summary tokens (2592) than image tokens (with 32 images, 6720), which acts as an information bottleneck: the model cannot afford one token per patch, so it must merge content that repeats across views into shared tokens.

At a middle layer of the LLM (layer 14 of 28), the hidden states at the summary-token positions are pulled out and passed through a small three-layer MLP head that decodes 32 3D Gaussians per token, producing 2592 × 32 Gaussians in total. Each Gaussian carries a 3D center, opacity, covariance, and spherical-harmonic color coefficients. These Gaussians are rendered at the original camera viewpoints with differentiable rasterization, and the rendered images are compared to the ground-truth frames with MSE and LPIPS losses — the photometric reconstruction loss.

Because learning Gaussians purely from photometric error is slow even without an LLM in the loop, the authors add a teacher: ZipSplat, a frozen feed-forward compact Gaussian estimator. They match ZipSplat's number of query tokens to their own number of summary tokens so the alignment is one-to-one, then add two distillation losses — one pulling the LLM's summary-token hidden states toward the teacher's refined query tokens via cosine similarity, and one matching the decoded Gaussian parameters with an L2 loss. The total objective is the language-modeling loss plus weighted reconstruction and distillation terms. Training uses the ScanNet-based datasets (SQA3D, ScanQA, Scan2Cap, ScanRefer, Multi3DRefer) plus a 100K subset sampled from SPAR-7M, on 16 GH200 GPUs with an effective batch size of 64 for 1 epoch. The teacher is only used during training.

Why This Matters

Impact on research. The paper challenges the dominant assumption that better 3D reasoning requires finer, pixel-level 3D supervision. It offers evidence that a coarse, object-centric reconstruction objective can propagate 3D-aware signals into an MLLM's own image features even though only a small set of dedicated tokens is directly supervised. It also connects cognitive accounts of human spatial perception to a concrete architectural design, and shows that reconstruction quality is not the goal — downstream reasoning is.

Real-world applications:

  • Robotics and embodied AI, where an agent must integrate observations from multiple camera views before acting.
  • AR/VR, where scene layout must be understood from a handful of headset or handheld views.
  • Assistive and household agents that answer questions such as "what is on the desk behind the chair" from a photo set of a room.
  • 3D content and inspection workflows, where a coarse but view-consistent scene representation from a few images can support querying and captioning.

Industry relevance. The results are achieved with a 7B backbone that outperforms much larger general-purpose models on SPAR-Bench, which is relevant for deployment cost. The method also avoids per-scene optimization and task-specific decoders, so one generalist model can handle question answering, captioning, and grounding through natural language.

Future Directions

  • Scaling and schedule. The paper's ablations stop at 4 epochs for the non-distilled variant and 1 epoch for the full model; how the method behaves under longer training or larger backbones is not established in the content provided.
  • Teacher dependence. Distillation from ZipSplat is what makes one-epoch training practical, and the teacher is frozen and training-time only. Whether the approach works without a compact Gaussian teacher, or with weaker teachers, remains open.
  • Bottleneck and token budget. Performance peaks at M = 2592 and drops at both 1296 and 5184, suggesting the trade-off between bottleneck strength and representational capacity is not fully understood, and that an adaptive or content-dependent number of summary tokens might work better.
  • Beyond the evaluated domains. All reported benchmarks are ScanNet-based or SPAR-Bench; generalization to outdoor, large-scale, or dynamic scenes is not reported, and the paper's own citations note that ScanNet-only training tends to overfit to ScanNet.

Target Audience

Researchers and graduate students working on multimodal large language models, 3D scene understanding, and spatial reasoning; practitioners who need agents that reason over multiple views of a scene; and readers interested in how cognitive accounts of human perception can be translated into model architectures. The paper is technical — it assumes comfort with transformer hidden states, cross-attention, Gaussian Splatting, and distillation — so beginners will get the most from the introduction, the analysis section, and the plain-language framing of the results.

Authors’ abstract

Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.

Read the original paper