Skip to content
AI.info

Research

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

Overview Research area: Evaluation of large language model (LLM) agents for 3D structure generation, spanning 3D asset synthesis, CAD assembly, part-based generation, and physical realizability. Techn

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures
arXiv
2610.04292
Published
2026-10-03
Authors
Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang, Sun Li, Jiayu Liu, Yifan Shen, Xu Cao, Jiarui Yao, Bingxuan Li, Ruhi Sarikaya, Heng Ji

AI summary

Overview

Research area: Evaluation of large language model (LLM) agents for 3D structure generation, spanning 3D asset synthesis, CAD assembly, part-based generation, and physical realizability.

Technical level: Advanced. The paper assumes familiarity with LLM agent tool-use loops, CAD/part assemblies, joints and kinematics, and 3D evaluation metrics.

Scope: LMBuild is a benchmark and interactive environment that measures whether LLM agents can generate 3D structures that are not only visually plausible but also structurally sound, functionally complete, and physically buildable and operable.

What This Paper Is About

Existing 3D-generation evaluations mostly score geometric quality, visual consistency, and semantic alignment, which says little about whether an object can actually be assembled and used. LMBuild reframes the problem as generating assembled structures — objects with meaningful part decompositions, joints, materials, and assembly sequences — and asks whether LLM agents can produce them from a label and a reference image. The paper builds an interactive construction environment, a curated benchmark, and a four-level evaluation framework, then tests 30 systems on it.

Key Contributions

  1. An interactive construction environment. Agents are given a standardized harness and a set of tools to retrieve components from a sub-part pool, create new parts with arbitrary geometry, place components at desired poses, and refine structures through inspection and revision. The environment enforces a maximum of 160 tool calls per episode, including at most 6 calls to check, 4 calls to render, and 80 placed parts, with part creation limited to 30 calls in Tier B and 60 calls in Tier C.

  2. A curated benchmark of buildable objects. From roughly 368K candidate 3D objects (367,991 total), the authors construct a full benchmark of 2,549 objects and a core set of 200 objects by repurposing established CAD datasets and augmenting them with real-world object knowledge extracted from Wikipedia, including core functional parts, object attributes, joints, and kinematic relationships. Average annotations are 7.62 parts / 9.25 affordances per core object and 7.03 parts / 8.83 affordances per full-benchmark object.

  3. A four-level hierarchical evaluation framework. The framework covers structural soundness, functional affordance, design quality, and physical realization, decomposed into twelve concrete sub-metrics (S.1–S.3, A.1–A.3, D.1–D.3, R.1–R.3) normalized to a common 0–100 scale.

  4. A 30-system empirical study. The evaluation spans 6 closed-source APIs, 13 open-source LLMs, and 11 domain-specific systems, plus ablations over part-access settings, revision rounds, and prompt detail.

Main Findings

  • Structural soundness and visual alignment are no longer primary bottlenecks. Frontier closed-source models achieve near-perfect soundness and strong design quality; for example, GPT-6 Astra scores 95.9 on collision-free structure (S.2) and 86.7 on structural stability (S.3), and Claude Fable 5.1 reaches 97.4 on S.2. Open-source models still show room for improvement in aesthetics and fine-grained visual alignment.

  • Functional affordance and physical operability remain the largest gaps. Joints and kinematics (A.3) are particularly hard: most open-source models score below 25, and even the strongest systems only slightly exceed 50. GPT-6 Astra reaches 63.0 on appropriate functional geometry (A.1) but only 45.3 on complete functional parts (A.2) and 41.7 on joints and kinematics (A.3). Simulation-based operability (R.3) is low across all systems; the paper describes best scores reaching only the mid-30s, while the reported table's highest R.3 value is 22.1 for Claude Opus 5.

  • Strong models create parts; weaker models retrieve. Under the three part-access settings, less capable models degrade sharply when required to create their own parts: Qwen3-VL-32B averages 61.7 across Soundness and Design under retrieval but drops to 34.1 under creation-only. Even adding creation alongside retrieval can hurt weaker models. Frontier models show the opposite trend, more often creating structure-specific components and maintaining comparable or stronger performance when retrieval is removed.

  • Explicit functional specifications help more than extra revision rounds. Three revision rounds yield overall gains of only about 2%. Enriching prompts with explicit attributes and functional requirements grounded in Wikipedia improves functional-part completeness by 10.3 points and physical operability by 12.0 points (p < 0.001), with additional improvements in kinematic correctness.

  • Frontier closed-source models converge. GPT-6 Astra, GPT-5.6 Sol, Claude Opus 5, and Claude Fable 5.1 exhibit broadly similar performance across many evaluation dimensions, suggesting leading model families are converging toward a comparable capability level for 3D structure generation.

  • Overall performance tracks general reasoning, but domain-specific systems stay competitive. Specialized systems occasionally surpass general models on individual metrics (e.g., Cube3D †(P) on D.2, Qwen3.5 models on R.1), suggesting targeted training and architectural specialization can partially compensate for weaker general reasoning.

  • A part-creation gap between open and closed models. Closed-source models succeed on 95.7% of part-creation attempts, versus 51.7% for open-source models, with Qwen3.5-27B at 89.4% as the only close exception. Failures cluster into two categories: violations of CAD-kernel conventions (such as listing polygon vertices clockwise) and invalid program structures (nesting nodes instead of referencing named nodes, using numeric node IDs, referring to undefined nodes).

  • Open models struggle with scene-state tracking. Open-source models refer to parts that were never placed at 7.2 errors per 100 calls versus 0.3 for closed-source models, and redundantly add existing entities (repeatedly placing parts instead of moving them, re-adding existing joints) at 3.2 errors per 100 calls versus 0.5 for closed-source models.

Methodology in Plain English

LMBuild treats construction as a conversation with a tool-using agent. Under the baseline protocol, the user supplies only a target structure label and a reference image — no geometric specifications, part-level instructions, or intermediate feedback — and the agent autonomously builds the object in a single interaction round, either retrieving parts from a candidate pool containing at least three times the required number of components or creating new parts when needed. It places and adjusts components, then specifies joints, materials, and an assembly sequence.

To isolate mechanisms, the authors run constrained variants: retrieval-only (agents must build entirely from the provided pool, Tier A), the baseline with retrieval plus creation (Tier B), and creation-only with retrieval disabled (Tier C). They also test multi-round revision with budgets of K ∈ 1, 2, 3 rounds, where the environment returns a six-view rendering and a rule-based diagnostic report covering collisions, disconnected components, stability, joint issues, and missing functional roles or materials. In a separate setting they augment the prompt with fine-grained descriptions of attributes and functional components.

For domain-specific systems that lack a common tool interface, the authors preserve native input and output protocols rather than adding unsupported capabilities; when a system produces geometry without articulation information, they apply Particulate to infer part-level articulation, joints, and kinematic constraints for evaluation. The benchmark data itself was built by stratifying across structural and kinematic diversity (six broad families spanning static load-bearing structures, hinged mechanisms, sliding mechanisms, rotating systems, serial kinematic chains, and multi-joint articulated mechanisms), requiring grounded objects whose properties could be verified from Wikipedia, diversifying part counts with at least four parts per object, and prioritizing product-level CAD from 27 open-source hardware projects. Each object category appears at most once in the core set.

Why This Matters

The paper argues that producing elegant geometry is fundamentally different from producing an object that can be built and perform its intended function, and shows that current evaluations overlook structural soundness, functional affordances, and physical realizability. By measuring where agents actually fail — functional parts, joints, kinematics, and simulated operability rather than appearance — LMBuild redirects research attention toward capabilities that matter for physical realization.

Real-world applications:

  • Robotic assembly of designed objects, where an agent's assembly sequence and part set must be executable rather than merely renderable.
  • Open-source hardware and product design, drawing on the 27 open-source hardware projects whose complete assemblies span robotics and vehicles, consumer devices, lab and medical instruments, and fabrication machines and tools.
  • Mechanical and engineering CAD, where the Fusion 360 Gallery assembly data from JoinABLe provides explicit part and joint structure.
  • Articulated everyday objects and furniture, where Artiverse contributes physically grounded objects with functional parts and kinematic annotations.

Industry relevance: The benchmark targets the gap between digital design tools and physical manufacturing. Models that specify materials, feasible assembly sequences, and operable joints could feed fabrication and assembly pipelines. The finding that prompt specification matters more than extra revision rounds also implies that product-side interfaces should elicit functional requirements from users before construction begins. The ethics statement notes performance on LMBuild should not be treated as sufficient evidence of readiness for safety-critical engineering or construction without domain-specific validation and human oversight.

Future Directions

  • Infer functional requirements before construction. Since explicit functional specifications produce the largest gains, the authors point toward agents that can infer, retrieve, or elicit an object's functional requirements themselves, reducing the need for users to supply detailed part-level and kinematic specifications.
  • Close the part-creation gap. Much of the gap appears to stem from weak adherence to tool-specific conventions and program structure rather than from deficiencies in geometric or spatial reasoning — an addressable target for training or better tooling.
  • Improve scene-state tracking in long-horizon construction. Open models repeatedly lose track of the evolving scene, referring to unplaced parts and re-adding existing entities, indicating a broader limitation in maintaining an accurate internal representation of the environment.
  • Extend the benchmark and validate in the physical world. All experiments here run on the 200-object core set while the 2,549-object full benchmark will be released; the conclusion notes that bridging benchmark results to real fabrication will require further validation involving manufacturing constraints, physical assembly, and functional testing.

Target Audience

Researchers and engineers working on LLM agents, 3D content generation, part-aware and articulated object synthesis, and CAD automation; benchmark developers interested in physical realizability and simulation-based evaluation; and teams evaluating whether model outputs can be manufactured or assembled, including robotics and open-source hardware communities.

Authors’ abstract

LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functional structures. LMBuild represents generated objects as assembled structures comprising part decompositions, joints, materials, and sequences. To support reproducible evaluation, we provide a unified framework consisting of: (1) an interactive environment in which agents can use tools to retrieve, create, and place components to construct objects; (2) a curated benchmark that repurposes established CAD datasets and augments them with knowledge from Wikipedia; and (3) a evaluation framework covering structural soundness, functional affordance, design quality, and physical realization. Evaluations across 30 systems reveal several intriguing findings: (a) Soundness and alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging; (b) stronger models more effectively create new components, whereas weaker models tend to rely on retrieval; and (c) providing functional specifications substantially improves part completeness, kinematics, and physical operability. These results show that generating real-world structures requires deeper reasoning about functional affordances, mechanics, and designing and creating novel components. We expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures.

Read the original paper