Research
EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling
Overview Research area: 3D computer vision, specifically benchmarking instruction-guided 3D asset editing and evaluating LLM/VLM agents that edit meshes by writing code. Technical level: Advanced. The

- arXiv
- 2610.02298
- Published
- 2026-10-01
- Authors
- Ruihan Yu, Yu-Ju Tsai, Muyao Niu, Runyi Li, Lian Fu, Hanqing Liu, Zheng-Hui Huang, Yonghao Yu, Sho Kuno, Ming-Hsuan Yang, Kaipeng Zhang, Zhixiang Wang
AI summary
Overview
Research area: 3D computer vision, specifically benchmarking instruction-guided 3D asset editing and evaluating LLM/VLM agents that edit meshes by writing code.
Technical level: Advanced. The paper defines new region-restricted voxel and image metrics, a deterministic data-synthesis engine, and a multi-method evaluation protocol; readers need familiarity with 3D representations, Chamfer distance, LPIPS/SSIM/PSNR, and agentic code-generation workflows.
Scope: EditHero introduces a 457-chain, 2755-edit benchmark with exact part-level ground truth after every turn, and uses it to compare 4 non-agentic 3D editing methods against 6 LLM/VLM agents under a self-rollout protocol.
What This Paper Is About
Existing text- and image-guided 3D editing methods are evaluated on a single edit applied to a clean source object, but real asset production in art and game pipelines is iterative: each turn changes part of an asset while preserving everything built before it — the same expectation behind the recent idea of "vibe modeling." This paper asks whether current methods can edit their own outputs repeatedly without losing earlier changes or degrading the asset, and builds EditHero, a benchmark of part-level edit chains with an exact 3D reference state after every turn, to answer that question.
Key Contributions
-
Generation engine. A data-synthesis engine that assembles library parts onto segmented host objects using four operations (add, remove, replace, retexture) and records every accepted operation in a JSON log, so replaying the log with the part library rebuilds the exact state after any turn.
-
Benchmark. EditHero, a human-reviewed benchmark for long-horizon part-level 3D editing containing 457 chains and 2755 edits across 252 hosts, with chains of 3 to 30 turns (median 6) and a median instruction length of 17 words.
-
Evaluation framework. A self-rollout protocol in which turn k takes the method's own output from turn k−1 as input, combined with region-restricted metrics that separate instruction following (IF) inside the edit region from content consistency (CC) outside it.
-
Empirical comparison. Evaluation of 4 non-agentic methods (PartFlow, Nano3D, 3DEditFormer, VoxHammer) and 6 LLM/VLM agents (GLM 5.3 Flash, DeepSeek V4.1 Flash, Sol 6, Astra, Fable 5.1, Opus 5.5) on both the full benchmark and a shared 55-chain (416-turn) comparison subset.
Main Findings
-
Whole-object scores reward doing nothing. On the whole object, the no-op baseline beats every non-agentic method on F-score, LPIPS and PSNR, because a single turn changes only a small part of the object. Only the edit region distinguishes real edits from inaction.
-
Non-agentic methods follow instructions poorly from the first turn. IF is 0.12–0.32 over all turns, and already 0.24–0.36 at turn 1, before any error can accumulate. Nano3D (0.31) and VoxHammer (0.32) follow instructions best among non-agentic methods, but VoxHammer keeps much less of the rest of the object (unchanged-region IoU 0.71 against Nano3D's 0.89).
-
Errors accumulate under self-rollout. From turn 1 to turn 7 on the 135 chains with at least 7 turns, CC⁰ falls from 0.60–0.78 to 0.33–0.67, with VoxHammer losing the most. In the reset diagnostic, starting each turn from the ground-truth previous state raises turn-5 IoU to 0.68 for Nano3D and 0.46 for 3DEditFormer, versus 0.52 and 0.36 under self-rollout — suggesting much of the loss is inherited rather than created in the current turn.
-
The regeneration baseline trades preservation for instruction following. TRELLIS regeneration scores the highest IF among non-agentic comparisons (0.50 on 457 chains, 0.49 on the 416-turn subset), but its unchanged-region IoU with its own previous output is only 0.38 (0.35 on the subset). No non-agentic method reaches regeneration on IF or the no-op baseline on preservation.
-
LLM/VLM agents preserve more and usually follow instructions better. On the 416-turn comparison subset, all 6 agents keep the rest of the object better than any non-agentic method (CC^prev 0.95–0.98 against at most 0.90), and all except GLM 5.3 Flash also follow instructions better (IF up to 0.62 for Opus 5.5, against at most 0.36).
-
Operation difficulty varies sharply for agents. Removal is easiest (IF 0.88–0.99) because code can delete a part outright. Additions and replacements are harder (IF 0.10–0.49 and 0.17–0.52), because new geometry built from code often misses the target's size or shape.
-
Every LLM beats the no-op baseline on whole-object LPIPS (0.136 for GLM 5.3 Flash down to 0.065 for Opus 5.5, against 0.154), and all of them deliver a mesh at every scheduled turn (416/416), whereas VoxHammer completes 377/416.
-
Agents are slow. An LLM needs about 1.5–6 minutes and 6–16 model calls per edit (medians), while a non-agentic method takes 20–50 seconds on a local GPU — too slow for interactive use.
-
Qualitative evidence on a covered wagon. Across 6 consecutive replacements, the non-agentic methods keep a blue or dark cover and end with dull olive instead of pale yellow, and 3DEditFormer and VoxHammer also change the color of the wagon bed, which no instruction touches. Opus 5.5, Astra and Fable 5.1 produce covers closer to the targets, and their errors stay on the cover.
Methodology in Plain English
The authors first build a data engine rather than collecting data by hand. A host object is divided into named slots, and candidate parts are drawn from a captioned library built from PartVerse-XL, Objaverse-XL and HY3D-Bench. For each turn, the engine picks one of four operations (add, remove, replace, retexture), retrieves a caption-matched part, screens candidates with a vision-language model, and checks placement for contact, scale, connectivity and interpenetration, looping until every check passes. Accepted operations are written to a JSON log; retexturing renders the slot and restyles the image with Qwen-Image-Edit, after which the texturing module of TRELLIS.2 textures the unchanged slot geometry. Because the log records every part, pose and texture, replaying it with the part library reconstructs the exact state after any turn. Human reviewers then check renders of every turn for wrong parts, orientation, placement, scale, ambiguous instructions and broken dependencies, and rewrite the drafted instructions.
For evaluation, each method runs in self-rollout: at turn k it receives its own output from turn k−1, the instruction and a target render, and ground-truth states are used only for scoring. All states in a chain share one coordinate frame and 4 fixed camera views (1 conditioning view, 3 held-out views). The edit region is defined from the parts named in the instruction before and after the turn — voxelized on a 64³ grid and dilated by one cell for 3D, or depth-tested and projected with one-pixel dilation for 2D — so lighting changes and retextures do not redefine where an edit is allowed. Two baselines anchor the scale: a no-op baseline that returns the initial object every turn (IF = 0, CC = 1), and a regeneration baseline that runs TRELLIS on the target render at each turn without the instruction or previous state. Metrics are computed separately in the edit region, the unchanged region and the whole object; instruction following is measured against the method's own previous output so that a part lost earlier does not count as a successful removal now.
Why This Matters
Research impact. The paper argues that whole-object F-score can be a poor indicator of editing ability, since returning the input already scores well, and supplies a benchmark whose targets are assembled deterministically rather than produced by a generative pipeline that may itself introduce unwanted changes. It is, to the authors' knowledge, the first benchmark to evaluate different natural-language part edits in sequence on one asset with an exact 3D reference after every turn, which reframes 3D editing evaluation from single-shot to sequential.
Real-world applications.
- Game and film asset production, where assets evolve through successive editing and revision rounds rather than a single authoring step.
- Iterative "vibe modeling" tools, in which users shape assets through a natural-language conversation and expect each turn to preserve prior work.
- Evaluating and selecting LLM/VLM agents for automated content pipelines, using measures that separately score the requested change and the rest of the asset.
- Diagnosing where multi-turn systems fail — inherited drift versus newly introduced error — via the reset diagnostic.
Industry relevance. The paper reports that agentic editing keeps untouched parts intact by construction but costs 1.5–6 minutes and 6–16 model calls per edit, against 20–50 seconds for non-agentic methods on a local GPU. That gap frames a concrete engineering target: an agentic and conventional combination that is both fast and accurate. The authors state they will release the data-synthesis engine and edit chains, as well as code at a public repository.
Future Directions
- Building geometry from code. Additions and replacements remain the hardest operations for agents (IF 0.10–0.49 and 0.17–0.52), because primitive-built geometry often misses the target's size or shape; improving this is an explicit open problem.
- Speed. An edit that takes minutes is too slow for interactive work, so reducing the 1.5–6 minutes and 6–16 model calls per edit is left open.
- Hybrid methods. The conclusion states it remains an open question whether a proper combination of agentic and conventional methods can be found that is both fast and accurate.
- Metric scope. IF is undefined when the relevant region is too small and for retextures, which require no geometric change; the paper reports how many turns have a defined IF, leaving extended coverage as a design question.
Target Audience
Researchers and engineers working on 3D generative models, 3D editing systems, and agentic content-creation pipelines; benchmark designers who need deterministic multi-turn ground truth; and practitioners in games, visual effects and interactive design tooling who want evidence on whether agentic or non-agentic editing better matches an artist's iterative workflow.
Authors’ abstract
3D editing methods are usually tested on a single edit, yet an asset is built through a long sequence of revisions, each of which must implement the requested change while leaving everything else unchanged. We introduce EditHero, to our knowledge the first benchmark for long-horizon, part-level 3D editing, with natural-language instructions and target images for both geometry and texture. A deterministic assembly engine produces the exact target after every edit, and every sequence is reviewed by hand. We use EditHero to compare 2 opposite approaches to 3D editing. Non-agentic methods operate top down, regenerating the object from a learned 3D representation and inferring what to keep. In contrast, LLM/VLM agents operate bottom up, editing through code that inspects the mesh and rewrites only the parts required by instructions. The non-agentic methods often miss the requested change and disturb regions that should stay fixed. Most LLMs follow instructions more closely, and all of them preserve the unedited parts better, but each of their edits takes minutes. We will release the engine and the edit sequences to support research on reliable iterative 3D editing.