Skip to content
AI.info

Research

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

Overview Research area: Software engineering and AI-agent evaluation, specifically benchmarks for LLM-based agents that generate executable Simulink models from natural-language engineering requiremen

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation
arXiv
2610.02304
Published
2026-10-01
Authors
Ruiqi Zhang, Jiahao Wang, Mingxuan Li, Haichen Luo, Chaoting Wang, Guoyu Mou, Keyu Lai, Hanchao Lv, Jiaxu Wang, Yibo Zheng, Aijun Yang, Xiaohua Wang

AI summary

Overview

Research area: Software engineering and AI-agent evaluation, specifically benchmarks for LLM-based agents that generate executable Simulink models from natural-language engineering requirements.

Technical level: Intermediate. The concepts (Simulink, .slx artifacts, native simulation, gating criteria, weighted scoring) are explained in the paper, but familiarity with model-based design and agent benchmarking helps.

Scope: The paper introduces SimuVerity, a 101-task, ten-domain benchmark that scores generated Simulink models against engineering requirements using native simulation rather than similarity to a reference model, and reports results from six agent systems.

What This Paper Is About

Existing Simulink benchmarks mostly ask whether a generated model compiles, runs, or looks like a reference model. The authors argue that neither question establishes whether a model actually satisfies its engineering requirements, since valid implementations can be structured differently from a reference and reference-aligned models can still be engineering failures. SimuVerity addresses this by grounding every task in an executable-system profile and scoring qualified models across six engineering dimensions through task-specific native simulation scenarios.

Key Contributions

  1. Benchmark curation. SimuVerity comprises 101 text-to-executable Simulink model-generation tasks across ten engineering domains. Each task includes a reconstructed or independently built reference system, an executable-system profile, and a profile-grounded engineering prompt.

  2. Evaluation framework. A hierarchical multidimensional framework combining three prerequisite gates (artifact delivery, native executability, engineering qualification) with four families of task-specific native simulation scenarios and six performance dimensions (A/Q/M/C/R/D).

  3. Empirical findings. Six agent systems were evaluated and their generated models analyzed from multiple engineering perspectives, including structural-similarity analysis, cross-domain capability gaps, visual layout quality, and tool ablations.

  4. Evaluator validation. Two domain experts independently rated 30 candidate models without seeing evaluator scores, yielding a weighted Cohen's kappa of 0.90 (95% CI [0.79, 0.96]) and a Spearman correlation between evaluator scores and mean expert ratings of 0.94 (95% CI [0.88, 0.98]).

Main Findings

  • Best agent scores only 42.86 overall. Claude Opus 4.8 with Claude Code reaches 42.86, followed by GPT-5.5 with Codex at 41.72, DeepSeek-V4-Pro with Claude Code at 28.40, Qwen3.8-Max with Claude Code at 25.01, GLM-5.3-Flash with Claude Code at 4.98, and Qwen3.8-27B-FP8 with Claude Code at 1.60. The 101 task-specific reference systems average 96.01, with 100.00% delivery, executability, and G-pass rates.

  • Gates do not imply engineering quality. Across all task–system runs, artifact delivery is 84.43%, native executability 68.65%, and G-pass 47.85%. For the top system, these figures are 94.72%, 86.47%, and 76.90%, meaning a substantial fraction of delivered, executable candidates are not qualified engineering implementations.

  • Output quality is the strongest dimension; robustness and dynamics are the weakest. Averaged equally across the six systems, end-to-end A/Q/M/C/R/D scores are 29.03, 38.46, 31.99, 29.21, 27.69, and 27.82, with operating-domain robustness (R) and dynamic response and recovery (D) lowest.

  • Structural similarity is a poor proxy for engineering performance. Across 95 delivered Opus Run 1 candidates, 39 were Reference-Aligned, 43 Partially Similar, and 13 Structurally Distinct; 56/95 (58.95%) were Partially Similar or Structurally Distinct. End-to-end means were 52.13, 44.05, and 42.37; post-G means were 59.80, 52.61, and 61.20. Of the 24 top-quartile candidates, 12 were Partially Similar or Structurally Distinct, while seven Reference-Aligned candidates scored zero, including two that passed the G gate.

  • Valid diversity and deceptive similarity coexist. Two Structurally Distinct candidates passed G and scored 85.10 and 80.07, while two Reference-Aligned candidates passed G but scored zero.

  • High scores do not imply readable diagrams. Of the 31 Opus Run 1 candidates with engineering scores of at least 70, experts judged eight severely disordered. Layout-only reorganization reduced block overlap to zero in every model, reduced line-crossing density by 22.7–85.1%, and reduced routing-detour ratio in seven cases, without changing any implementation.

  • Simulation feedback and MCP tools are critical. On ten cross-domain, high-performing tasks, Full MCP produced 100.00% delivery, 100.00% executability, 100.00% G pass, and an overall score of 78.46. No-Simulation MCP dropped these to 53.33%, 33.33%, 13.33%, and 7.88, with lower scores on all ten tasks. Batch-only dropped them to 60.00%, 50.00%, 40.00%, and 18.23, with 8 of 10 tasks scoring lower.

  • Capability gaps differ by domain. For Claude Opus 4.8 with Claude Code, the authors identify qualification-dominated domains (power electronics and fluid systems), post-qualification bottlenecks (battery and aerospace systems), and mixed bottlenecks (mechanical and communication systems).

Methodology in Plain English

The authors built a benchmark from the requirements side rather than the artifact side.

  1. Collect and reconstruct reference systems. Candidate systems were drawn from public examples, engineering projects, research artifacts, and expert-built models, then retained if they could be executed in Simulink, showed meaningful engineering behavior, and broadened coverage. Each was reconstructed into a reproducible evaluation target by resolving dependencies, exposing the signals needed for control and observation, and connecting it to repeatable operating conditions and a task-specific evaluator.

  2. Profile each system. Domain experts executed each reference system under a unified protocol and wrote an executable-system profile capturing system boundary, interfaces, mechanisms, control relationships, operating conditions, and dynamic behavior.

  3. Write prompts from the profiles. A unified prompt framework covers the modeled system, deliverable, input–output interfaces, operating range, functional objectives, required mechanisms, control relationships, and dynamic behavior. Prompts were checked against the profile for consistency and native realizability.

  4. Build simulation tests. Each task's requirements were translated into an executable test suite spanning four families: operating-envelope scenarios, interaction and fault scenarios, temporal-process scenarios, and mechanism and causal-check scenarios. In total there are 595 formal scenarios: 205 operating-envelope, 189 interaction and fault, 79 temporal-process, and 122 mechanism and causal-check.

  5. Gate, then score. Candidates must first deliver a complete .slx artifact, then complete native loading, diagram updating, compilation, and simulation, then pass an engineering-qualification gate checking required functional roles, connected signal paths, and valid evaluation evidence. Qualified candidates are scored on six dimensions: A (objective accuracy), Q (output quality), M (mechanistic fidelity), C (control and causal integrity), R (operating-domain robustness), and D (dynamic response and recovery). Dimension applicability varies by task: A 101, Q 101, M 93, C 35, R 101, D 98.

  6. Aggregate. Observed quantities are converted to utilities in [0,1] via task-specific functions, aggregated into dimension scores on a 0–100 scale, and combined with expert-set task weights through a weighted geometric mean so a strong dimension cannot fully offset a critical weakness.

  7. Run the experiments. Each task–system pair was run three times from a fresh session and directory with a separate run identifier, and the final artifact of each run was evaluated with the same frozen task-specific scorer. Runs begin in a clean session with the same core MATLAB MCP tools, and the execution environment hides the MATLAB example-model directories, local reference-model directories, and scenario and scoring-script directories from the agent. Non-delivery, technical non-execution, and candidate-caused runtime failures count as failures; API, evaluator, or environment errors count as infrastructure failures and are rerun.

Example: for the electric-axle task XAUTO-07, one Opus Run 1 candidate received A 100.00, Q 100.00, M 78.25, C 71.88, R 99.90, and D 100.00, with expert weights 0.25, 0.10, 0.25, 0.20, 0.10, and 0.10, producing a task score of 88.03.

Why This Matters

SimuVerity reframes how generated engineering artifacts should be judged: not by whether they resemble a reference implementation, but by whether they meet requirements across prescribed operating conditions. It also shows that requirement-level evaluation surfaces failures — qualification failures, robustness failures, dynamic-response failures, and readability failures — that similarity-based metrics hide. The authors note that the benchmark relies on MATLAB/Simulink and that native-simulation evaluation can be slow.

Real-world applications:

  • Automotive and aerospace control development, where generated models must hold up under boundary conditions, payload variation, and fault scenarios rather than just compile (the benchmark includes tasks such as a closed-loop tiltrotor VTOL flight model, an electric-aircraft mission and powertrain model, and an electric axle drive with regenerative braking).
  • Power electronics, robotics, thermal, fluid, battery, communication, and biomedical system modeling, where the benchmark's ten domains map to distinct engineering practice areas.
  • Toolchain and agent platform evaluation, where the ablations show that simulation feedback and MCP-style native tool access materially change outcomes on tasks the agent can otherwise complete.
  • Model review and maintenance workflows, since layout-only reorganization reduced block overlap to zero and cut line-crossing density by 22.7–85.1% without touching implementation, and since reference-aligned models were sometimes the weakest performers.

Industry relevance: the results show that current agents cannot be trusted to produce engineering-grade Simulink models end-to-end. Practitioners would need a closed loop of specification-driven modeling, native simulation validation, and multi-condition feedback correction, and reviewers would still need to inspect requirement satisfaction directly rather than accept structural resemblance.

Future Directions

  1. Optimize the evaluation pipeline for efficiency and scalability. The authors explicitly state this as future work, citing the reliance on MATLAB/Simulink and the slowness of native simulation-based evaluation.

  2. Close the qualification-to-performance gap. G-pass rates (47.85% across all runs) far below delivery rates (84.43%) show that forming genuine engineering implementations remains a bottleneck distinct from satisfying multidimensional requirements after qualification.

  3. Improve multidimensional performance after qualification. End-to-end and post-G results show operating-domain robustness (R) and dynamic response and recovery (D) are the weakest dimensions, raising the question of what training or feedback would raise them.

  4. Address diagram readability alongside engineering accuracy. Eight of 31 high-scoring candidates were severely disordered, so how to make agents produce organized, traceable diagrams without changing model semantics remains open.

Target Audience

Researchers building or evaluating agentic engineering tools; benchmark designers working on executable-artifact generation; Simulink and Model-Based Design practitioners interested in where LLM agents currently fail; and evaluators of code- or model-generation agents who want evidence that similarity-to-reference metrics are insufficient. The validation study and per-dimension breakdowns also make the paper useful to teams designing requirement-level scoring rubrics.

Authors’ abstract

Existing Simulink benchmarks mainly evaluate whether generated models compile, execute, or resemble a reference model. These criteria do not establish whether a model satisfies its engineering requirements. We introduce SimuVerity, a benchmark of 101 text-to-executable Simulink model-generation tasks across ten engineering domains. For each task, executable-system profiles ground the engineering specification and four families of native simulation scenarios. A hierarchical evaluator first checks artifact delivery, native executability, and engineering qualification, then scores qualified models across six dimensions covering accuracy, output quality, mechanistic fidelity, control and causal integrity, operating-domain robustness, and dynamic response. We evaluate six agent systems with SimuVerity. The best system achieves an overall score of only 42.86. The results show that structural similarity is a poor proxy for engineering performance: capability bottlenecks arise both in producing qualified implementations and in satisfying multidimensional requirements after qualification. Meanwhile, some high-scoring models still exhibit severe visual-layout disorder. SimuVerity provides a systematic basis for assessing agents' engineering capabilities and diagnosing failures in executable Simulink model generation.

Read the original paper