Research
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments? Overview Research area: Artificial Intelligence — evaluation benchmarks for autonomous computer-use agents operati

- arXiv
- 2609.37686
- Published
- 2026-09-29
- Authors
- Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun, Haoyang Li, Yipeng Wei, Naihao Xue, Xiaohan Yu, Zhuo Tao, Yihe Zang, Yajiao Wang, Jingyi Tang, Yi Li, Jingjing Zhou, Jie Luo, Bohan Zeng, Chengyu Shen, Hao Jiang, Chong Chen, Bowen Qu, Olive Huang, Zeqiang Wang
AI summary
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?Overview
Research area: Artificial Intelligence — evaluation benchmarks for autonomous computer-use agents operating professional industrial engineering software.
Technical level: Advanced. The paper assumes familiarity with agent benchmarks, GUI/CLI interaction paradigms, and domain-specific artifact formats (STEP files, netlists, G-code, IFC models).
Scope: This paper introduces EngiWorld, a benchmark of 1,301 expert-curated tasks across 6 engineering domains and 26 software platforms, together with an artifact-centric verifier suite, and uses it to measure how seven frontier models perform on professional engineering workflows.
What This Paper Is About
Autonomous agents have become capable at general computer use, but the paper argues that reliable automation of professional engineering is still out of reach because engineering workflows require reasoning over geometric and physical constraints that must be preserved across many software and design stages. Existing benchmarks mostly judge success by comparing final states or reference solutions, cannot verify whether an engineering artifact is actually functional, and provide almost no native domain feedback such as rendered output, solver convergence curves, or design-rule violation reports. EngiWorld is built to evaluate the complete design loop rather than a single link in it.
Key Contributions
-
The first benchmark structured around the complete design loop. EngiWorld contains 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms and workbenches, with both GUI and CLI interfaces and 6 task types ranging from software selection to fully open-ended work.
-
An artifact-centric evaluation methodology. A unified domain-verifier suite programmatically checks geometric validity, physical feasibility, and rule compliance of both final and intermediate artifacts (such as STEP files, netlists, and G-code), instead of relying on final-state comparison or reference solutions.
-
Continuous scoring for quantitative design. Rather than binary success, quantitative design tasks are gated by feasibility checks and then scored continuously by how well the specification is attained, evaluating how well a task is completed rather than merely whether it is completed.
-
A broad empirical evaluation with diagnostic ablations. Seven frontier models are evaluated, and ablations vary initial-image presentation, runtime image access, accessibility information, screenshot resolution, and interaction-history length across three models (GPT-5.6 Sol, Gemini 3.7 Flash, and Kimi K3).
Main Findings
-
Frontier models fall far short of industrial standards. The strongest model, Claude Opus 5, achieves an overall EngiScore of only 44.3, followed by GPT-5.6 Sol at 38.0, Qwen3.8 Max at 25.8, DeepSeek V4.1 Flash at 25.5, Gemini 3.7 Flash at 25.3, Kimi K3 at 24.1, and Qwen3.8 Flash at 16.1. All seven models receive zero scores on 128 of the 300 evaluated tasks.
-
Cross-application work collapses. Only 6 of 168 Multi-software attempts succeed across all models, corresponding to the reported 3.6% success rate for multi-software attempts. Claude Opus 5 reaches 8.3 on the Multi-software category score, while Gemini 3.7 Flash, Kimi K3, and Qwen3.8 Max score 0.0 there.
-
Single-tool operation is much easier than coordination. Claude Opus 5 completes roughly half of the Single-software tasks (50.3) and two thirds of the Image-based modeling tasks (66.7), but solves only one quarter of the Software-selection tasks (25.0). The paper concludes that agents can execute individual stages but break down at transitions between them.
-
Interface matters, and CLI strength does not transfer to GUI. Claude Opus 5 scores 47.9 on CLI and 40.5 on GUI; GPT-5.6 Sol scores 47.1 on CLI but 28.6 on GUI. DeepSeek V4.1 Flash achieves 48.3 on CLI — the highest CLI score outside Claude and GPT — yet succeeds on only three GUI tasks, giving it a GUI score of 2.0.
-
Domains show model specialization. Claude leads in CAD, CAE, EDA, and BIM and ties GPT in CAM. Qwen3.8 Max performs best in 3D visualization despite its lower overall score. CAM is the most difficult domain for every model.
-
Cost and effort vary widely. Models average 54.9–120.5 interaction steps per task; GPT-5.6 Sol uses the fewest (54.9) and Claude Opus 5 next (70.2). DeepSeek V4.1 Flash and Gemini 3.7 Flash score similarly at approximately 25, but DeepSeek costs less than half as much — a gap the paper attributes to pricing rather than frugality, since DeepSeek generates 9.9K output tokens per step against Gemini's 0.7K. Claude reaches the highest score at nearly twice GPT's API cost (19.02 vs 9.73 dollars per task).
-
Two failure modes dominate. Declared completion with a zero score and decision-turn exhaustion together account for 94.1% of failures. Declared completion is the dominant failure for GPT (87.0% of retained failures) and Claude (73.9%), versus 39.6% for Gemini. Decision-turn exhaustion accounts for 55.3% of GUI failures and 36.6% of CLI failures. Runtime-limit failures are concentrated in long-running executions, including 28 DeepSeek runs and 9 Claude runs that reach the five-hour limit.
-
Observation design helps unevenly. Presenting reference drawings in the task message raises GPT-5.6 Sol's GUI score from 15.0 to 31.7 and Gemini 3.7 Flash's from 3.3 to 10.0, but for Kimi K3 it moves the CLI score down from 40.0 to 36.7. Adding an accessibility tree improves Kimi, changes GPT little, and lowers Gemini's score, while increasing API cost by 50.7–147.9%. The native 1920×1080 resolution produces the highest score for all three ablated models; a 0.6× scale cuts GPT's cost by 30.8% while preserving most of its score but degrades Kimi much more.
-
Optimal interaction history is model dependent. Gemini benefits from extending the window from five to fifteen turns, GPT largely plateaus after ten, and Kimi performs best at ten — for Kimi, extending from ten to fifteen turns lowers EngiScore while increasing cost by 53.0%.
Methodology in Plain English
Each task is defined by a tuple of components: a goal, task inputs (reference media, initial project files), a software environment and workspace state, a set of permitted interface actions, required deliverables, and acceptance criteria. The agent interacts with a real native engineering application in an isolated Windows or Ubuntu virtual machine through either a GUI (via screenshots and mouse or keyboard actions) or a CLI (terminal commands and scripts), and the episode ends when the agent declares DONE or FAIL or hits a task-specific decision limit. Crucially, evaluation ignores the action trajectory and the final screen state: instead, verifiers reopen the submitted artifacts and extract properties such as geometric dimensions and topology, simulation quantities, schematic connectivity, building-model entities and relationships, and 3D scene structure. Binary tasks require all acceptance criteria to pass; quantitative design tasks must first pass a feasibility gate before receiving a continuous quality score, and infeasible submissions are scored zero.
Tasks were built through an expert-led, AI-assisted pipeline: experts selected representative real-world workflows from documentation, tutorials, and textbooks; AI helped standardize task specifications and implement verifiers; experts then built a validated reference realization in the actual software and reviewed each full task package. Tasks with identified inconsistencies were revised and revalidated. The evaluation ran on a stratified subset of 300 tasks (152 CLI, 148 GUI) allocating four or five tasks to each of 73 software-and-category strata, covering all six domains. Agents received the most recent 15 interaction rounds, with default decision limits of 200 rounds for GUI and 100 for CLI, raised to 300 and 150 for Multi-software and Open-ended tasks.
Why This Matters
The paper exposes a measurable gap between general computer-use competence and the demands of professional engineering: the strongest model tested reaches an EngiScore of 44.3, and cross-application workflows remain nearly unsolved. Because the benchmark verifies artifacts rather than screens, it distinguishes agents that look finished from agents that produce usable engineering output — and it shows that agents frequently declare completion without verifying deliverables, and overlook feedback such as solver residuals.
Real-world applications implied by the covered domains:
- Computer-aided design and manufacturing — producing valid solid models and machining files across tools such as FreeCAD, SolidWorks, AutoCAD, and SolidCAM.
- Simulation and analysis — running and interpreting multiphysics and CFD workflows in ANSYS, Abaqus, OpenFOAM, FEniCS, CalculiX, FLORIS, and OpenFAST.
- Electronic design automation — working with schematics and board designs in Cadence OrCAD, EAGLE, Altium Designer, and KiCad.
- Building information modeling and 3D visualization — editing and coordinating building models and scenes in Archicad, Revit, Bonsai, OpenStudio, Blender, and ZBrush.
Industry relevance: the 1,723 task–software associations and the 37 environment configurations (26 software-specific, 10 composite for multi-software workflows, and one dedicated open-ended environment) reflect the breadth of real engineering toolchains, making the benchmark a plausible yardstick for vendors and teams building agents intended for professional production use rather than general desktop automation.
Future Directions
- Closing the design loop rather than individual stages. The paper's central diagnostic is that agents execute single stages but break at transitions; how to preserve parametric logic and intermediate-artifact dependencies across applications remains open.
- Making agents verify before declaring done. Declared completion was the dominant failure mode for GPT and Claude, so teaching agents to self-check artifacts against engineering constraints is a direct target.
- Exploiting native engineering feedback. Agents frequently overlook signals such as solver residuals; how to incorporate rendered outputs, convergence curves, and design-rule reports into the decision policy is unresolved.
- Model-dependent interface and observation design. Because the benefit of accessibility trees, extra visual evidence, and longer history varied by model — and sometimes reversed or reduced scores — there is no single best configuration; how to adapt observation design per agent and per task is an open question.
Target Audience
Researchers and engineers working on autonomous agents, computer-use benchmarks, and AI for engineering automation; benchmark designers who need artifact-level verification rather than final-state comparison; and practitioners in CAD, CAE, CAM, BIM, EDA, and 3D visualization evaluating whether current agents can be trusted with professional toolchains. Readers seeking a beginner-level introduction to agent evaluation should expect a steep learning curve, since the paper presumes background in agent interaction models and engineering artifact formats.
Authors’ abstract
Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.