Skip to content
AI.info

Research

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

LEGO-Anything: Coding Agents for 3D Scene Reconstruction Overview Research area: Computer vision and 3D scene reconstruction, specifically single-image scene reconstruction using general-purpose codin

LEGO-Anything: Coding Agents for 3D Scene Reconstruction
arXiv
2609.36380
Published
2026-09-28
Authors
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

AI summary

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Overview

Research area: Computer vision and 3D scene reconstruction, specifically single-image scene reconstruction using general-purpose coding agents that write executable Blender programs, plus a benchmark for evaluating them and a study of whether the resulting scenes serve as image representations.

Technical level: Intermediate. The paper is readable without deep graphics expertise, but assumes familiarity with 3D scene representations, rendering, and the coding-agent (tool-use plus execution feedback) paradigm.

Scope: The paper introduces an Image-to-Code framework for reconstructing 3D scenes from one RGB image, a simulator-grounded benchmark (LEGO-Bench), a training-free harness plugin (LEGO-Plugin), and a downstream evaluation (LEGO-World) that queries reconstructed scenes for detection, segmentation, and depth.

What This Paper Is About

Reconstructing a 3D scene from a single image is most useful, the authors argue, when the output is an explicit executable program rather than a fixed mesh, point map, or object set, because a program can be run, inspected, edited, and queried. Existing pipelines either compose specialized modules for perception and asset retrieval or map images directly to fixed 3D outputs, and neither gives a direct way to check the result against the input image and revise it. The paper studies what happens when a general-purpose coding agent performs this reconstruction by iteratively writing and executing Blender code, comparing renderings to the reference image, and revising the scene program.

Key Contributions

  1. LEGO-Anything, an Image-to-Code (Image2Code) framework in which a coding agent alternates between editing Blender code, executing it, and inspecting the resulting scene and renderings, producing a construction trajectory of intermediate programs, scenes, and observations whose final state is the submitted scene program.

  2. LEGO-Bench, a simulator-grounded benchmark with 208 RGB inputs rendered from 104 scenes across 8 environments and 17 themes, using 443 registered assets, built in LychSim from Fab scene and asset packs. It separately scores artifact Validity, visible-surface Reconstruction, and rendered Appearance, and is designed to be extensible while retaining precise automatic evaluation.

  3. A failure diagnosis plus LEGO-Plugin, a training-free harness plugin (exposed as MCP tools, workflow skills, and runtime hooks) with three modules — Enhanced Initialization, Grounded Refinement, and Version Control — each targeting one recurring failure observed in agent construction trajectories.

  4. LEGO-World, an evaluation setting where object detections, instance masks, and relative depth are obtained as deterministic readouts of a single frozen reconstructed scene, tested on 100 images each from COCO val2017, LVIS v1 val, and ETH3D.

Main Findings

  • Coding agents deliver valid artifacts reliably, but fidelity lags. All six GPT configurations achieve near-saturated Validity, yet overall scores range from 53.4% indoor / 39.6% outdoor for GPT-6-astra down to 15.4% / 15.3% for the strongest GPT-5.6 configurations (reported as 15.4% indoor for GPT-5.6-terra and 15.3% outdoor for GPT-5.6-sol).

  • GPT-6-astra leads overall. With the Codex harness, GPT-6-astra scores 100.0 Validity, 52.4 Reconstruction, 54.4 Appearance, and 53.4 overall indoors, and 98.0 / 34.0 / 45.5 / 39.6 outdoors. Gen3DSR attains higher Reconstruction indoors than GPT-6-astra (65.4% vs. 52.4%) but produces no evaluable appearance (0.0), while VIGA and 3D-RE-GEN deliver valid artifacts with low fidelity.

  • Baselines are narrower in coverage. Three of six baselines do not support outdoor scenes, whereas all coding agents run on both splits with stable Validity. SceneConductor reaches only 0.3 overall indoors (12.3 Validity), REST3D 2.8, SceneGen 3.4 indoor and 0.5 outdoor, 3D-RE-GEN 6.3 indoors.

  • More complexity degrades fidelity, not executability. Averaged over all GPT-6 and GPT-5.6 configurations, Validity stays at 99.5% across tiers while the overall score drops from 24.6% (Easy) to 21.3% (Medium) to 20.5% (Hard). The steepest drop is outdoor Reconstruction, from 18.4% to 12.7%. Outdoor scenes are harder than indoor scenes at every tier.

  • More reasoning effort helps, mostly for stronger models. On a fixed 42-case Office subset, GPT-6-astra rises from 32.3% to 61.8%, GPT-6-sol from 21.3% to 39.7%, and GPT-6-luna from 14.4% to 21.2% as reasoning effort goes from Low to XHigh. GPT-5.6 models show weak or non-monotonic changes.

  • Construction quality is non-monotonic. GPT-6-astra reaches its first evaluable scene within roughly the first tenth of its budget, while GPT-5.6-sol needs about one fifth. For GPT-5.6-sol, 29.6% of edits decrease the score, and its final submission trails its best intermediate scene by 3.2 points.

  • Agents cannot judge their own scenes. Across all 36 builder–judge pairs of six models, self-judgments agree 45.8% of the time on Reconstruction and 62.2% on Appearance, versus 45.4% and 63.9% for cross-model judgments (chance is 50%). Self-judging beats the mean of the other five judges for only 3 of 6 builders on Reconstruction and 2 of 6 on Appearance.

  • LEGO-Plugin helps every model, most for the weakest. On the 42-case Office subset, the three GPT-5.6 models improve by 55–63% relative (about 7–8 points), GPT-6-luna by 27.5%, GPT-6-sol by 12.1%, and GPT-6-astra by only 2.1%, with up to 62.7% relative gains in overall score reported in the abstract.

  • Reconstructed scenes support vision tasks but trail specialists. With GPT-6-astra scenes (no plugin), readouts reach 30.14 box AP on COCO val2017 versus 59.88 for DINO, 14.75 mask AP on LVIS v1 val versus 53.96 for Segment Anything 3, and 0.1554 AbsRel on ETH3D versus 0.0783 for Depth Anything 3 — roughly half of specialist box AP, without any task-specific training.

  • Cost and quality trade off across the model family. GPT-5.6-luna is the absolute cheapest configuration, but most of the quality–cost frontier is defined by the GPT-6 family, with GPT-5.6-sol and GPT-5.6-terra lying off the frontier.

Methodology in Plain English

The researchers reformulate single-image 3D reconstruction as a coding task. Instead of asking a model to predict a scene in one shot, they let a general-purpose coding agent interact with Blender: it writes code, runs it, looks at the rendered result and the evolving scene, compares that to the input photograph, and edits the code again. The final submission is a scene program plus an export and a rendered view.

To measure how well this works, they build a benchmark from professionally authored simulator scenes rather than photographs, because that gives realistic-looking inputs while keeping exact ground truth private to the evaluator — geometry, depth, instance masks, object correspondences, and scene transforms. Scenes are grouped into themes with nested Easy, Medium, and Hard object sets so that architecture, materials, lighting, and camera stay fixed while content grows, enabling matched comparisons. Scoring is split into three axes: whether the submission is a usable artifact at all (can the file open, is the export well formed, is the image parseable), how well visible surfaces match the reference geometry in camera coordinates using object-level F1 with a depth-scaled tolerance of 5% of forward depth, and how visually close an evaluator re-render is to the reference, measured as the fraction of pixels within an error threshold of 30 on a 0–255 scale.

To understand failures, they re-score every renderable intermediate artifact in the agent trajectories, measuring how quickly an agent produces its first evaluable scene, how often edits make things worse, and how far the final submission is from the best intermediate state. They also run a judging experiment where models pick the better of two renderings and are checked against the deterministic metric direction. These diagnoses motivate a plugin that stabilizes initialization, replaces self-judgment with measurement against the reference image, and rolls back harmful edits. Finally, they freeze scenes reconstructed from natural images and treat detection, segmentation, and depth as deterministic queries on those scenes.

Why This Matters

Impact on research. The paper reframes 3D reconstruction as program synthesis with an inspectable, editable, queryable artifact, and it introduces a diagnostic benchmark that separates "did the agent produce something valid" from "is the scene geometrically and visually faithful." It also provides an empirical diagnosis of three concrete agent failure modes (weak initialization, regressive edits, unreliable self-evaluation) that generalize beyond 3D graphics to any iterative coding-agent loop with a perceptual objective.

Real-world applications:

  • Content creation for games, AR/VR, and visual effects, where an editable Blender scene derived from a single reference image is more useful than a fixed mesh.
  • Simulation and robotics, since an executable scene program can in principle be run, scaled to new environments and difficulty levels, and reused rather than treated as a static asset.
  • Image editing and compositing workflows that need explicit object state, camera control, and manipulable scene structure.
  • Downstream vision pipelines, since the paper shows a single frozen scene can yield detections, masks, and relative depth without task-specific training.

Industry relevance. The evaluations are run in a shared execution environment based on the Harbor Framework with Blender 5.0.1 and Blender-MCP, under the Codex harness, and the plugin is delivered as MCP tools, skills, and runtime hooks — an integration pattern directly compatible with production agent harnesses. The cost–performance analysis (token-cost proxy versus benchmark score) speaks to deployment tradeoffs, and the finding that a training-free plugin substitutes for some of what weaker models lack is relevant to teams choosing between a stronger model and a cheaper one plus scaffolding.

Future Directions

  • Closing the gap between valid artifacts and faithful reconstruction: Validity is near-saturated while Reconstruction and Appearance remain limited, especially outdoors, and the paper reports that weaker agents regress after reaching an evaluable scene.
  • Improving agent self-evaluation, since geometric judgment sits at or below chance (45.8% self, 45.4% cross-model against a 50% baseline), which the authors argue requires grounding refinement in deterministic evidence rather than self-assessment.
  • Making reconstructed scenes precise enough to serve as representations of natural images, given that readouts reach 30.14 box AP versus DINO's 59.88, 14.75 mask AP versus Segment Anything 3's 53.96, and 0.1554 AbsRel versus Depth Anything 3's 0.0783. The paper notes that scene exports provide no calibrated confidences, so box and mask AP use an equal-confidence protocol, leaving calibrated confidence as an open problem.
  • Extending the benchmark itself: the paper states that users can convert their own scenes and assets into new cases through the same construction pipeline, and it includes an NYC aerial split without difficulty pairing as a large-scale reconstruction stress test, suggesting more environments, difficulty levels, and views as natural next steps.

Target Audience

Researchers and practitioners in 3D computer vision, scene reconstruction, and generative graphics; agent researchers studying coding agents, tool use, harness design, and iterative self-correction; benchmark designers interested in simulator-grounded evaluation that separates artifact validity from fidelity; and engineers building production pipelines that need editable, queryable 3D representations rather than fixed outputs.

Authors’ abstract

A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.

Read the original paper