Skip to content
AI.info

Research

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

Overview Research area: Multimodal large language model (MLLM) spatial reasoning, combined with on-policy self-distillation (OPSD) and tool-using coding agents. Technical level: Advanced. The paper as

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
arXiv
2610.11366
Published
2026-10-08
Authors
Rongxue Li, Meng Yang, Yiru Mao, Yongliang Tao, Lulu Hu, Bin Yang, Zhao Xu, Weihua Luo, Bowen Xu

AI summary

Overview

  • Research area: Multimodal large language model (MLLM) spatial reasoning, combined with on-policy self-distillation (OPSD) and tool-using coding agents.
  • Technical level: Advanced. The paper assumes familiarity with KL-divergence distillation objectives, unlikelihood training, policy-gradient methods such as GRPO, and agentic code-execution loops.
  • Scope: The paper asks whether a 9B-parameter MLLM can internalize the spatial reasoning behavior of a tool-using spatial coding agent so that it operates with no tools at inference time, and proposes SpatialOPSD plus Repetition-Aware Distillation to do so.

What This Paper Is About

Spatial reasoning requires inferring latent 3D properties from 2D observations, and the intermediate geometric steps are hard to annotate or verify, which limits standard supervised fine-tuning and reinforcement learning. Recent spatial coding agents sidestep this by writing and executing code with geometry tools, but they pay a large inference-time latency and infrastructure cost and depend on external tools. The paper's goal is to transfer the reasoning behavior of these verified agents into a standalone MLLM via on-policy self-distillation, so that the final model reasons spatially in a single tool-free forward pass.

Key Contributions

  1. An observation that agent traces elicit spatial reasoning. The authors show that simply conditioning chain-of-thought prompting on a spatial coding agent's execution trace matches the accuracy of the full tool-using agent, and increases references to 3D perception concepts (1.8 times) and counterfactual motion concepts (1.6 times) relative to generic CoT prompting.
  2. SpatialOPSD, an on-policy self-distillation framework that treats verified spatial coding agent traces as privileged information. A frozen copy of the student's initial checkpoint is conditioned on summarized traces and supervises the student, which sees only the raw input.
  3. Repetition-Aware Distillation, which combines a token-level mask that excludes self-repetitive and privileged-content-transcribing spans from the distillation loss with an unlikelihood penalty applied to those same flagged tokens.
  4. A demonstration of a self-distillation loop for geometric reasoning, in which a model generates its own verified agent traces and then trains on them, with gains replicated across three student scales (2B, 4B, 27B).

Main Findings

  • Agent-trace conditioning is nearly as strong as the agent itself. Across four inference settings (Direct QA, Prompted CoT, Spatial Coding Agent, Agent-Trace-Conditioned CoT) evaluated with Qwen3.5-9B, trace-conditioned CoT matched the full agent's spatial reasoning capability without further tool calls.
  • SpatialOPSD beats SFT and GRPO on average. On MindCube-Tiny and ViewSpatial, SpatialOPSD-9B reaches a reported average of 51.9, versus 45.9 for the Qwen3.5-9B-CoT baseline, 47.6 for SFT, and 48.2 for GRPO.
  • Component-level results differ from the average. On the MindCube-Tiny Overall column specifically, SFT scores 58.1 and GRPO 53.3, while SpatialOPSD scores 56.9. On the ViewSpatial Overall column, SpatialOPSD scores 46.8 against 37.0 for SFT and 43.1 for GRPO. The average advantage comes largely from avoiding SFT's collapse on ViewSpatial.
  • Tool-free performance approaches the tool-using teacher. SpatialOPSD retains 91% of the performance of its iterative, tool-dependent teacher SpatialClaw (51.9 vs. 57.3), and exceeds GPT-5 (51.0) on the same average while remaining below Gemini-3-Pro (60.7).
  • Repetition-Aware Distillation trades a little accuracy for large leakage reduction. Vanilla OPSD shows an 82.7% privileged-information leakage rate. Mask alone gives 53.0 average accuracy with 73.6% leakage; unlikelihood alone gives 51.5 accuracy with 6.7% leakage; combining both gives 51.9 accuracy with 9.1% leakage.
  • Hyperparameters matter asymmetrically. With L_min fixed at 32, raising λ_rep from 0.00 to 0.15 lowers leakage from 73.6% to 0.8% while MindCube-Tiny accuracy moves from 60.0 down to 55.1. Varying L_min from 4 to 64 changes accuracy little but swings leakage between 75.9% and 9.1%. The chosen settings are n = 4, L_min = 32, λ_rep = 0.1, ε = 10⁻⁶.
  • The "summary" privileged-information variant works best. It reaches 56.9 on MindCube-Tiny Overall and 46.8 on ViewSpatial Overall, ahead of fact (55.2 / 45.4), full-trace (54.5 / 45.8), and intent (53.3 / 47.4). The authors attribute this to a moderate teacher–student entropy gap.
  • Verification of traces matters more than trace quantity. Training on all 10,000 MindCube training samples without ground-truth verification yields 52.0 on MindCube-Tiny Overall and 46.3 on ViewSpatial; training on only the 5,008 verified traces yields 56.9 and 46.8, an improvement of 4.9 and 0.5 points respectively.
  • The method scales across model sizes. Average accuracy rises from 40.6 to 42.8 at 2B, from 46.4 to 50.2 at 4B, and from 50.6 to 54.2 at 27B, corresponding to gains of 2.2, 3.8, and 3.6 points.
  • Out-of-distribution performance is largely preserved. On OmniSpatial, BLINK

Authors’ abstract

Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM can internalize this agentic capability to operate entirely tool-free. We begin with a simple observation: prompting an MLLM with summarized execution traces of a spatial coding agent naturally unlocks the model's internal spatial Chain-of-Thought (CoT). Motivated by this, we introduce SpatialOPSD, an on-policy self-distillation framework that internalizes spatial reasoning into a standalone MLLM by formulating verified agent traces as privileged information. To mitigate privileged-information leakage during distillation, we introduce Repetition-Aware Distillation, which combines repetition masking with unlikelihood regularization. Experiments across multiple benchmarks demonstrate that self-distilling SpatialOPSD achieves higher average accuracy than SFT and GRPO on both spatial and OOD datasets, exhibiting superior performance and generalization.

Read the original paper