Skip to content
AI.info

Research

SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning

Overview Research area: Computer Vision, specifically multi-view spatial reasoning with vision-language models (VLMs), 3D reconstruction, and chain-of-thought (CoT) supervision. Technical level: Advan

SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
arXiv
2609.33616
Published
2026-09-27
Authors
Yang Cao, Jiaxin Zhang, Dave Zhenyu Chen, Yingji Zhong, Ruiyuan Gao, Lanqing Hong, Dan Xu

AI summary

Overview

Research area: Computer Vision, specifically multi-view spatial reasoning with vision-language models (VLMs), 3D reconstruction, and chain-of-thought (CoT) supervision.

Technical level: Advanced.

Scope: The paper introduces SpatialSpeak, a two-stage training framework that first teaches a VLM multi-view 3D reconstruction through question-answering targets and then trains it to reason explicitly over the learned geometry to answer spatial questions.

What This Paper Is About

Multi-view VLMs are increasingly given 3D geometric priors from pretrained reconstruction models, but the usual answer-only training never directly supervises the intermediate geometric estimates a model needs for quantitative spatial questions. The authors hypothesize that spatial chain-of-thought supervision works better when the model has first jointly learned complementary local geometry (where a marked image point sits in 3D) and global scene context (where object instances sit across views). SpatialSpeak connects these two stages, QA-Native Reconstruction Pretraining (QA-RP) and spatial Chain-of-Thought with Visual Compensation (CoT-VC), so that geometry is learned and then explicitly used within the same autoregressive text interface.

Key Contributions

  1. Demonstrating that reconstruction pretraining amplifies reasoning supervision. The paper establishes that the gain from CoT-VC rises from 2.6 points without QA-RP to 6.9 points with QA-RP on ReVSI, and reports state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench.

  2. QA-Native Reconstruction Pretraining (QA-RP). A Stage I method that learns local point geometry and global object-layout context through text-based QA targets expressed in a shared 3D coordinate system (the first-frame camera coordinate system), with no auxiliary 3D regression heads or extra geometric losses.

  3. Spatial Chain-of-Thought with Visual Compensation (CoT-VC). A Stage II method that explicitly supervises question-relevant geometric estimates, task-specific answer derivations, a reliability assessment (High or Low), and visually grounded answer refinement when the geometric estimate is judged unreliable.

  4. Ablations isolating each design choice, covering local versus global reconstruction supervision, the reliability threshold tau, and the separate contributions of spatial CoT and visual compensation.

Main Findings

  • Overall performance: SpatialSpeak-4B reaches a ReVSI average score of 62.8, exceeding the strongest compared baseline, SpatialStack-4B (54.1), by 8.7 points. It obtains the best reported result on five of the seven ReVSI task categories, including object counting (64.9 vs. 36.3), room size (66.0 vs. 54.4), and relative distance (71.9 vs. 62.9).

  • Component contributions: On ReVSI, the variant with neither component scores 52.4; QA-RP alone reaches 55.9 and CoT-VC alone reaches 55.0; combining both gives 62.8, an improvement of 10.4 points over the baseline. Removing QA-RP or CoT-VC from the full method reduces the score by 7.8 or 6.9 points, respectively.

  • Local and global reconstruction supervision are both useful: The full method scores 62.8, versus 59.9 when global object-center queries are removed, 59.3 when local point queries are removed, and 55.0 when both are removed.

  • Reliability threshold matters: For the relative-error threshold tau used to build reliability labels, tau = 0.3 achieves the highest score of 62.8, compared with 59.6 for tau = 0.1 and 59.5 for tau = 0.5. All three settings outperform the variant without CoT-VC (55.9).

  • Both CoT and visual compensation help: Starting from 55.9 without CoT-VC, adding spatial CoT without VC improves the score to 58.5; incorporating VC raises it further to 62.8, an additional gain of 4.3 points.

  • Metric-scale reconstruction quality: On 200 randomly sampled ScanNet validation scenes held out from training, SpatialSpeak's Stage I model achieves the lowest errors without alignment (Acc. 8.9 cm, Comp. 9.3 cm), compared with CUT3R (13.7 and 12.7 cm) and MapAnything (36.3 and 28.4 cm). After GT Sim(3) alignment it obtains 5.0 cm on both metrics, lower than MapAnything (5.7 and 6.2 cm) but slightly higher than CUT3R (4.7 and 4.6 cm).

  • VSI-Bench (normal training setting): SpatialSpeak averages 63.3, the highest among compared methods, exceeding Omni-View-7B (55.4) by 7.9 points.

  • VSI-Bench (scaled training setting): SpatialSpeak reaches 73.0 with a 4B backbone, compared with 72.6 for GeoThinker-8B.

  • SPAR-Bench: SpatialSpeak achieves an average score of 76.0, exceeding SpatialStack-4B (72.0) and GeoThinker-8B (68.2) by 4.0 and 7.8 points, respectively.

  • Qualitative behavior: In one ReVSI example, asked how many blackboards are in the scene, SpatialSpeak enumerates two blackboard instances with their estimated 3D centers and predicts a count of 2, matching ground truth, whereas the variant with neither QA-RP nor CoT-VC predicts 3. A further example shows it enumerating four chair instances and predicting 4 versus 6 for that same variant, and a reconstruction example estimates a marked distance as 1.24 m against a ground truth of 1.25 m.

Methodology in Plain English

The authors start from a baseline built on VG-LLM that takes N RGB frames plus a natural-language query and produces a text answer. Visual patch tokens come from a vision encoder, while a frozen VGGT multi-view geometry encoder processes all frames jointly to produce 3D-aware features that are fused with the visual tokens. Only the LLM parameters are trained; the vision encoder, geometry encoder, and MLP projector stay frozen.

Stage I, QA-RP, turns reconstruction into ordinary question answering. A local query marks a 2D point in one frame with a red cross and asks the model to output that point's 3D coordinates in the first-frame camera coordinate system; ground truth comes from ScanNet depth maps and known camera extrinsics. A global query asks the model to detect the 3D center points of objects across all frames in the first-frame coordinate system, producing a list of semantic labels and 3D centers, ordered by first appearance and then by spatial location. Both are trained with the standard next-token prediction loss on a mixture of local and global samples, so geometry learning shares the same output interface used later for reasoning.

Stage II, CoT-VC, builds structured responses from ScanNet spatial QA pairs enriched with 3D object annotations. Deterministic CoT generators cover four question types: distance, size, count, and closest object. For each, the system derives a 3D estimate from the annotations and compares it against the ground-truth answer to assign a reliability label using a relative-error threshold tau (count uses exact integer equality; closest object uses letter match). Each training response contains a 3D reasoning chain followed by a reliability token: if the estimate is reliable it outputs High plus the 3D-derived answer; otherwise it outputs Low, a visual refinement note, and the ground-truth answer. This trains the model to admit when its geometry is approximate and fall back to direct visual inspection. Question types not covered by these templates, such as relative direction, retain direct answer supervision to preserve coverage of the full VSI-Bench task distribution.

Implementation uses Qwen3-VL-4B as the backbone. All stages share a cosine learning rate schedule with a warm-up ratio of 0.03, weight decay of 0.01, and 32 frames sampled per video clip. Stage I trains for 1 epoch with a global batch size of 32 and a peak learning rate of 5e-6; Stage II trains for 1 epoch with a global batch size of 64 and a peak learning rate of 1e-5. Stage II training data follows VG-LLM's use of subsets from the LLaVA-Hound split of LLaVA-Video-178K and SPAR-7M, augmented with the constructed spatial CoT data.

Why This Matters

Impact on research. The paper reframes geometry injection into VLMs: rather than only fusing features from a reconstruction model, it makes the VLM itself perform reconstruction through QA and then supervises how that geometry is used. The measured interaction, where QA-RP roughly triples the benefit of CoT-VC on ReVSI (from 2.6 to 6.9 points), gives a concrete argument that geometric pretraining and reasoning supervision are complementary rather than interchangeable, and the reliability-and-refinement mechanism offers a template for handling approximate intermediate estimates in other quantitative reasoning settings.

Real-world applications:

  • Embodied agents and service robots that must judge object counts, distances, room sizes, and spatial relations from a stream of camera frames before acting.
  • Augmented and virtual reality, where placing or aligning virtual content requires metric-scale geometry and reliable answers about object layout.
  • Navigation and route-planning assistants that need to reason about relative distances and directions across multiple viewpoints.
  • Inspection, mapping, and 3D content-creation workflows that need reconstructed scene geometry plus natural-language querying of that geometry.

Industry relevance. The framework reaches its reported scores with a 4B backbone, and Stage I does not add regression heads or extra geometric losses, so the design is comparatively lightweight relative to larger spatial baselines. This matters for teams deploying multimodal models in robotics, AR/VR, and spatial analytics, where the ability to audit an intermediate geometric estimate and decide when to trust it is as valuable as the final answer.

Future Directions

  • Extending CoT templates beyond the four covered question types. Relative direction and other tasks are handled with direct answer supervision rather than geometric CoT; designing deterministic generators for them is an open step.
  • Closing the alignment gap in Stage I. Under GT Sim(3) alignment, SpatialSpeak's reconstruction (5.0 cm on both metrics) is slightly behind CUT3R (4.7 and 4.6 cm), even though it leads without alignment, so the source of that reversal is unresolved.
  • Scaling the backbone. Results are reported for a 4B model; whether the QA-RP and CoT-VC interaction persists or grows at larger parameter counts is not reported.
  • Generalizing beyond the current data and geometry encoder. Training and evaluation center on ScanNet, LLaVA-Video-178K subsets, SPAR-7M, and the frozen VGGT encoder; behavior on other scene domains or with different geometry providers is not reported.

Target Audience

Researchers and engineers working on multimodal large language models, 3D scene understanding, and spatial reasoning, particularly those building embodied agents, robotics perception, or AR/VR systems. It is also relevant to readers interested in chain-of-thought supervision and in how models can learn to assess and correct their own intermediate estimates. The paper assumes familiarity with vision-language model architectures, multi-view reconstruction, and standard spatial reasoning benchmarks.

Authors’ abstract

Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction. We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.

Read the original paper