Research
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
Overview Research area: Benchmarking large language models (LLMs) on scientific code generation, specifically the finite element method (FEM), matrix structural analysis (MSA), and related computation

- arXiv
- 2512.20732
- Published
- 2025-12-23
- Authors
- Saeed Mohammadzadeh, Erfan Hamdi, Joel Shor, Emma Lejeune
AI summary
Overview
Research area: Benchmarking large language models (LLMs) on scientific code generation, specifically the finite element method (FEM), matrix structural analysis (MSA), and related computational mechanics tasks.
Technical level: Intermediate. Readers benefit from basic familiarity with LLM evaluation concepts and with the broad idea of numerical simulation / finite element analysis, but the paper does not require deep numerical-methods expertise to follow its argument.
Scope in one sentence: The paper introduces FEM-Bench, a small, carefully curated benchmark of introductory but nontrivial computational mechanics coding and unit-testing tasks, and reports baseline results for several leading LLMs on it.
What This Paper Is About
Existing code-generation benchmarks (such as HumanEval, MBPP, SWE-Bench, and DS-1000) test general programming logic, software engineering, or tool use, and mathematics benchmarks test symbolic or quantitative reasoning, but none evaluate whether an LLM can carry out the physical modeling, numerical discretization, and structured computational reasoning needed to write correct physics-based simulation code. FEM-Bench addresses this gap by posing self-contained computational mechanics tasks, each requiring a numerically correct Python function and a set of pytest unit tests that can distinguish correct implementations from known-incorrect ones. The goal is a diagnostic challenge suite rather than a large-scale training dataset, prioritizing interpretability and reasoning depth over task count.
Key Contributions
- The FEM-Bench framework, including its design principles and scope, built around self-contained Python task modules with reference implementations, dependency helper functions, pytest test functions, known-failing examples, and standardized
task_info()metadata. - A suite of tasks and evaluation tools grounded in canonical computational mechanics concepts, covering FEM and Matrix Structural Analysis, with tiered scaffolding (T1/T2/T3) that modulates difficulty by controlling how many helper functions are provided.
- A dual evaluation protocol in which code synthesis and unit-test synthesis are treated as intertwined, first-class components: generated functions are compared against reference outputs on curated verification inputs, and generated tests must pass on the reference implementation while failing on bundled expected-failure implementations.
- A baseline evaluation of leading LLMs, characterizing current performance and failure modes across the benchmark's 2025 release.
Main Findings
- Function-writing leader: In a five attempt run, the best performing model at function writing, Gemini 3 Pro, completed 30/33 tasks at least one out of five times, and 26/33 tasks five out of five times. (The figure 33 is the denominator reported in these scores.)
- Unit-test-writing leader: The best performing model at unit test writing, GPT-5, had an Average Joint Success Rate of 73.8%.
- Broad spread across models: Other popular models showed a broad range of performance on the benchmark; the paper does not report individual numeric scores for those models in the provided content.
- State-of-the-art models do not reliably solve the suite: Despite the tasks being introductory and aligned with a first graduate course on FEM, no model solved all of them reliably. A model that succeeds only intermittently on a task is a different reliability profile from one that never succeeds, and the benchmark reports both.
- Even the simplest task is non-trivial: The 12x12 local elastic stiffness matrix for a 3D Euler-Bernoulli beam is described as one of the simplest tasks in FEM-Bench 2025, yet its accompanying tests probe symmetry, rigid-body modes, block-level consistency of axial, torsional, and bending terms, and agreement with closed-form Euler-Bernoulli beam theory, so the task demands more than formula transcription.
- Four distinct capabilities are probed: domain knowledge, compositional reasoning, algorithmic fidelity, and self-verification. Domain knowledge is evaluated by tasks such as T3 variants that withhold helper functions and standalone tasks such as
MSA_3D_local_geometric_stiffness; compositional reasoning by the T1/T2/T3 tiering and tasks such asMSA_3D_elastic_critical_load; algorithmic fidelity by reference-output matching on curated verification inputs; self-verification by joint test success rate over expected-failure cases. - Failures are interpretable: Because computational mechanics has a mature culture of verification, validation, and uncertainty quantification, incorrect outputs can be traced to programming logic, geometry handling, numerical implementation, or floating-point-sensitive operations rather than treated as an undifferentiated error.
Methodology in Plain English
Each benchmark task is a single Python module. It contains a reference implementation of one well-defined numerical or physical quantity with a prescribed function signature and a detailed docstring (for example, computing a local element stiffness matrix, evaluating shape functions and their derivatives, assembling a global stiffness matrix, or running a full analysis). It also contains pytest-style test functions with descriptive docstrings, optional dependency helper functions, and known-incorrect implementations that serve as expected failures. A single task_info() function packages all of this into a structured dictionary.
The FEM-Bench software loads these tasks, extracts and normalizes the source code into a Task object, and automatically builds two standardized prompts per task: one asking a model to write the function, and one asking it to write tests. Prompts are saved to disk before inference for reproducibility. Outputs are parsed, validated, and evaluated: generated code is run against reference outputs on curated verification inputs, and generated tests are run both against the reference implementation (they must pass) and against the bundled known-incorrect implementations (they must fail). Results are aggregated into per-model and per-task scores.
Difficulty is modulated through tiering. For content whose typical implementation spans many helper routines, such as elastic critical load analysis (which requires a linear elastic solve, geometric stiffness assembly, and a generalized eigenvalue solve), the authors define separate tasks with different amounts of provided scaffolding, isolating whether a failure comes from a specific sub-component or from composing the components. All tasks in FEM-Bench 2025 were manually written by the authors, with LLM assistance limited to narrow sub-function contributions that were reviewed and verified; each task began from a working codebase developed via test-driven development and validated against multiple analytical solutions.
Why This Matters
Impact on research. FEM-Bench supplies a benchmark where correctness criteria are objective and quantitatively measurable, and where failures can be attributed to specific kinds of reasoning. This makes it possible to track progress on physics-grounded code generation separately from general programming ability, and to study the coupling between writing code and writing the tests that verify it. It also positions computational mechanics as a testbed for broader questions about physical reasoning and world modeling in AI.
Real-world applications. The paper identifies domains that depend on physics-based simulation grounded in computational mechanics:
- Robotics.
- Digital twins of aircraft and human systems.
- Climate modeling.
- Engineering design and optimization.
Industry relevance. Simulation code that runs but produces a non-convergent, unstable, or physically impossible result has no scientific value, so correctness is non-negotiable in these workflows. Companies and national laboratories that rely on simulation software have a direct interest in whether LLM-generated analysis code can be trusted, and the benchmark's tiering and error attribution give a practical way to see where models need supervision versus where they can be relied upon. The paper also notes that fine-tuning open-weight models has been explored for finite element code (Deotale et al., 2026, using FEniCS) and for course-specific expert models integrated into an AI-University educational platform (Shojaei et al., 2025), and that agent-based benchmarks such as FEABench primarily measure navigation of professional simulation software APIs rather than the underlying physical and numerical reasoning.
Future Directions
- Increasing task sophistication: The authors state that future iterations of FEM-Bench will incorporate increasingly sophisticated tasks to track progress as models evolve, extending beyond material aligned with a first graduate course.
- Reliability rather than best-of-n success: The gap between succeeding at least once out of five attempts (30/33 for Gemini 3 Pro) and succeeding five out of five times (26/33) raises the question of how to raise consistency, not just peak capability, on scientific code.
- Improving self-verification: With the best unit test performance at an Average Joint Success Rate of 73.8%, the ability to write discriminative, physics-aware tests that fail on incorrect implementations remains a clear open problem.
- Separating API navigation from physical reasoning: The paper argues that benchmarks which mainly test use of professional simulation software APIs do not decouple documentation interpretation from implementation of physical and numerical reasoning; building tasks that isolate each is an open direction.
- Fine-tuning and education: The cited work on fine-tuning open-weight models for finite element code and course-specific expert models suggests follow-up questions about whether targeted training closes the gaps FEM-Bench exposes.
Target Audience
Researchers evaluating LLM reasoning and code generation, especially those working on scientific machine learning and AI for science; computational mechanics and finite element researchers and educators who want a concrete way to assess AI assistance in their domain; and engineers or tool developers building LLM-assisted simulation and analysis workflows who need to know where current models are and are not dependable. Readers looking for a large-scale training corpus will not find one here, since the benchmark is deliberately a small diagnostic suite.
Authors’ abstract
As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to generate scientifically valid physical models has become a critical gap. Computational mechanics, which develops and applies mathematical models and numerical methods to predict the behavior of physical systems under forces, deformation, and constraints, provides an ideal foundation for structured scientific reasoning evaluation. Problems follow clear mathematical structure, enforce strict physical and numerical constraints, and support objective verification. The discipline requires constructing explicit models of physical systems and reasoning about geometry, spatial relationships, and material behavior, connecting directly to emerging AI goals in physical reasoning and world modeling. We introduce FEM-Bench, a computational mechanics benchmark designed to evaluate the ability of LLMs to generate correct finite element method (FEM) and related code. FEM-Bench 2025 contains a suite of introductory but nontrivial tasks aligned with material from a first graduate course on computational mechanics. These tasks capture essential numerical and physical modeling challenges while representing only a small fraction of the complexity present in the discipline. Despite their simplicity, state-of-the-art LLMs do not reliably solve all of them. In a five attempt run, the best performing model at function writing, Gemini 3 Pro, completed 30/33 tasks at least once and 26/33 tasks all five times. The best performing model at unit test writing, GPT-5, had an Average Joint Success Rate of 73.8%. Other popular models showed broad performance variation. FEM-Bench establishes a structured foundation for evaluating AI-generated scientific code, and future iterations will incorporate increasingly sophisticated tasks to track progress as models evolve.