Skip to content
AI.info

Research

GENIUS: An Agentic AI Framework for Autonomous Design and Execution of Simulation Protocols

Overview Research area: Agentic AI for computational materials science — specifically, using large language models plus a structured knowledge graph to automatically write, validate, and debug Quantum

GENIUS: An Agentic AI Framework for Autonomous Design and Execution of Simulation Protocols
arXiv
2512.06404
Published
2025-12-06
Authors
Mohammad Soleymanibrojeni, Roland Aydin, Diego Guedes-Sobrinho, Alexandre C. Dias, Maurício J. Piotrowski, Wolfgang Wenzel, Celso Ricardo Caldeira Rêgo

AI summary

Overview

Research area: Agentic AI for computational materials science — specifically, using large language models plus a structured knowledge graph to automatically write, validate, and debug Quantum ESPRESSO density functional theory (DFT) input files.

Technical level: Intermediate. The framing and motivation are accessible to a general reader, but evaluating the results requires some familiarity with DFT input decks (namelists, cards, k-points, pseudopotentials) and with LLM agent architectures.

Scope: The paper describes the GENIUS framework — a recommendation system, a protocol generation module, and an automated error-handling loop — and benchmarks it on 295 human-written calculation prompts, reporting an overall success ratio of 235/295 ≈ 0.7966.

What This Paper Is About

State-of-the-art electronic-structure codes like Quantum ESPRESSO are accurate and freely available, but configuring them still demands specialist knowledge of syntax, parameter interdependencies, and error messages. This "know-do gap" excludes experimentalists and slows Integrated Computational Materials Engineering (ICME), where the tools are mature but the human interface is the bottleneck.

The goal of GENIUS is to close that gap: let a researcher type a free-form, natural-language request for a simulation, and have an AI system autonomously produce a valid, runnable Quantum ESPRESSO input file — repairing its own failures when the code crashes — without the user touching the code's documentation.

Key Contributions

  1. A smart Quantum ESPRESSO knowledge graph (KG). The KG encodes the pw.x documentation as 247 nodes and 330 connectivity edges, extracted from a plain-text version of the online docs and converted to key-value form using anthropic/claude-3.5-sonnet, followed by manual verification. It captures parameter names, connections, and logical conditions (including implicit ones, e.g. a "Cu surface" request automatically activates the "Metallic systems" condition).

  2. A three-part agentic architecture under a finite-state machine. The framework combines (i) a recommendation system, (ii) a protocol generation module, and (iii) an automated error handling (AEH) system, orchestrated by a finite-state machine with an Entry Node, State Transitions, Error Recovery States, and Terminal States (Success/Failure).

  3. A model-agnostic, tiered LLM hierarchy with automatic escalation. Protocol generation uses two worker models (databricks/dbrx-instruct, meta-llama/llama-3.1-405b-instruct) and one referee model (anthropic/claude-3.5-sonnet), with support tasks handled by mistralai/mixtral-8x22b-instruct and google/gemini-2.0-flash-001. Each model gets three attempts before the system switches to the next, more capable model.

  4. A benchmark of 295 human-authored prompts with a documented complexity rubric and semantic analysis. Prompts were collected from chemists and physicists who routinely run DFT but not Quantum ESPRESSO, and complexity was scored using an LLM-based rubric inspired by Brown et al.

Main Findings

  • Overall success is roughly 80%. Of the 295 calculation requests, 235 produced a valid simulation protocol, giving P(S) = 235/295 ≈ 0.7966.

  • Zero-shot success is 14.24%. 42 requests reached a valid protocol on the very first attempt without invoking error handling — P(ZS) = 42/295 ≈ 0.1424. Broken down by complexity, that is 9.4% basic, 7.2% standard, and 1.3% complex (17.9% of all runs in the paper's alternative framing).

  • Most failures are repaired autonomously. The abstract reports that 76% of cases are autonomously repaired, and that success decays exponentially with attempt number down to a 7% baseline once Model 1's initial attempts are exhausted.

  • Prompt complexity is not the limiting factor. Successful runs at each attempt mix basic, standard, and complex prompts. The authors argue complex prompts can even help, because they carry more distinctive instructions. The prompt set scored 44.3% basic, 48.5% standard, and 7.2% complex.

  • The framework, not the strongest model, drives success. The referee model (anthropic/claude-3.5-sonnet) is not used disproportionately — a plateau in success appears at a baseline level rather than a spike at the top of the model hierarchy. The authors read this as evidence the framework is model-agnostic and that its architectural intelligence, not raw LLM power, does the work.

  • LLM-only baselines fail almost entirely. Base LLMs used without the GENIUS framework contributed negligibly to producing valid QE input files with correct cards and mutually consistent parameters, regardless of nominal reasoning enhancements — attributed to the models lacking embedded crystallographic information and the ability to infer QE's keyword/card interdependencies.

  • Inference cost halves and hallucinations nearly vanish. Compared with LLM-only baselines, GENIUS halves inference costs and "virtually eliminates hallucinations."

  • The prompt space forms five main semantic clusters. Of 295 prompts embedded as 3072-dimensional vectors with OpenAI's text-embedding-3-large model and mapped onto a 10×10 self-organizing map, five high-activation neurons were well separated, plus residual prompts. The SOM's Topological Error was 0.0373 and its Quantization Error 0.4970; inputs were unit-normalized (maximum possible pairwise distance 2.0). Manual inspection split prompts mainly into structural relaxation and single-shot DFT calculations.

  • A worked example self-heals in about three minutes. For a user prompt requesting a geometry optimization of 2D PdS₂ in the P2₁/c space group with the B3LYP functional, GENIUS generated a valid input specifying 20% exact exchange, a plane-wave basis, smearing, a mixing parameter, a 7×7×2 k-point mesh, BFGS relaxation, and Pd and S pseudopotentials. The timeline shows one QE crash, one retry within the AEH loop, then steady execution and completion in approximately 3 min.

  • Validation is execution-based. Quantum ESPRESSO PWSCF v.7.2 is the execution-level validator: a protocol is valid only if the simulation finishes with no CRASH file and a zero exit code.

Methodology in Plain English

The team built a pipeline that sits between a researcher's typed request and the Quantum ESPRESSO program.

First, they turned the QE pw.x documentation into a structured graph. Because QE organizes inputs into nameless sections and cards with different syntax, they converted the online HTML into plain text, extracted the two types separately, and used anthropic/claude-3.5-sonnet to produce key-value pairs, which they then checked by hand. The result is a graph of 247 nodes and 330 edges that a user can browse interactively.

Second, a recommendation system reads the user's prompt. It extracts material information (routing to the Materials Cloud MC2D or MC3D database depending on whether the target is two- or three-dimensional, with an agent that standardizes the retrieved geometry for reproducibility) and pulls relevant QE parameters from the graph using three combined routes: nodes that are always required, keyword search over the raw node text (top 70% by cosine similarity, using a hashing vectorizer with 2^17 features), and QE-specific conditions. The authors report extracting 162 conditions across nine categories in one passage and 164 unique conditions across nine categories in another. Each matched condition activates connected nodes; the union is then evaluated parameter by parameter, and irrelevant parameters get None and are dropped. The output is a template with suggested values and data types (CHARACTER, REAL, INTEGER, LOGICAL) plus atomic positions and pseudopotentials.

Third, LLMs write the actual input file from that template. The paper uses two prompt-engineering strategies: one that provides contextual scaffolding and asks the model to reason step by step (critical for error handling), and one that extracts structured JSON using explicit schemas and few-shot examples, parsed back with regular expressions into a Python dictionary.

Fourth, the generated file is run. If QE crashes, a CRASH file describes the error. An LLM extracts keywords from it, queries the knowledge graph, and proposes a fix. Each model gets three retries; logs from failed attempts within the same retry cycle are dropped to save context. If a model exhausts its retries, the system switches to the next model in the hierarchy and restarts from the recommendation system's original template — the new model sees only the error message, the relevant documentation, the latest protocol, and the original prompt. If everything fails, the workflow terminates with a failure message.

The whole thing runs as a FastAPI RESTful service. A POST /workflow/ endpoint accepts a JSON payload with the calculation prompt, model hierarchy, and per-task model configuration, alongside endpoints for status (GET /workflow-status/{workflow_id}), results (GET /results/{workflow_id}), and timeline visualization (GET /timeline/{workflow_id}), plus real-time Server-Sent Events at /logs. Supporting libraries include ASE, scikit-learn, and networkx.

Why This Matters

Impact on research. The paper reframes the problem as one of implementation science rather than algorithmic accuracy: DFT codes already agree closely with each other and approach experimental precision, so the bottleneck is the human interface. By automating protocol generation, validation, and repair, GENIUS aims to improve reproducibility, re-usability, and transferability of simulation protocols, and to widen who can participate in computational materials research. The authors position it as both an ICME accelerator and a test of whether agentic AI can handle precise scientific tasks where off-the-shelf LLMs fail.

Real-world applications:

  • Battery materials screening — automated setup for large batches of electrode or electrolyte structures.
  • Catalyst design — running relaxations and single-shot calculations for surface and adsorption systems without manual input-deck construction.
  • Structural alloys — bulk metallic systems, where the KG auto-invokes the "Metallic systems" condition from a prompt.
  • 2D materials — the demonstrated example is a 2D PdS₂ geometry optimization routed to the MC2D database.

Industry relevance. The framework is model-agnostic and provider-agnostic, so an organization can swap in cheaper, self-hosted, or newer LLMs without changing the architecture — and could self-host models to mitigate the LLM API latency the authors identify as the main secondary constraint on turnaround time. Halving inference cost versus LLM-only baselines matters directly for anyone running this at scale. The REST API and web dashboard design means it can be embedded in existing computational workflows or used point-and-click by non-programmers.

Future Directions

  • Improving the knowledge graph's connections and conditions. The authors explicitly flag this as future work requiring specialized expertise, and invite community contribution — suggesting the current 330 edges do not fully capture QE's parameter interdependencies.
  • Extending beyond Quantum ESPRESSO. The paper states the approach can be extended to any atomistic simulation code, including molecular dynamics, but only QE is demonstrated here.
  • Reducing latency from LLM API providers. The authors note that the long parameter-evaluation time is an external constraint; self-hosting would allow parameters to be processed in parallel. The paper does not report wall-clock throughput across the full 295-prompt benchmark.
  • Understanding what the residual failures are. Success decays to a 7% baseline after Model 1's attempts, but the paper does not characterize which error classes remain unsolved or how the AeH-only conditional success probability breaks down (that formula is cut off in the content available).

Target Audience

This paper is most useful to computational materials scientists and ICME practitioners who want to automate DFT workflows; to researchers building domain-specific LLM agents who want a concrete example of grounding a model in a curated knowledge graph with a finite-state error-recovery loop; to tool developers at national labs, HPC centers, and software companies building interfaces to simulation codes; and to implementation-science and research-software-engineering audiences interested in the "know-do gap" framing of why mature tools go unused. Beginners can follow the motivation and headline results, but the methodology assumes comfort with DFT input structure and agent design patterns.

Authors’ abstract

Predictive atomistic simulations have propelled materials discovery, yet routine setup and debugging still demand computer specialists. This know-how gap limits Integrated Computational Materials Engineering (ICME), where state-of-the-art codes exist but remain cumbersome for non-experts. We address this bottleneck with GENIUS, an AI-agentic workflow that fuses a smart Quantum ESPRESSO knowledge graph with a tiered hierarchy of large language models supervised by a finite-state error-recovery machine. Here we show that GENIUS translates free-form human-generated prompts into validated input files that run to completion on $\approx$80% of 295 diverse benchmarks, where 76% are autonomously repaired, with success decaying exponentially to a 7% baseline. Compared with LLM-only baselines, GENIUS halves inference costs and virtually eliminates hallucinations. The framework democratizes electronic-structure DFT simulations by intelligently automating protocol generation, validation, and repair, opening large-scale screening and accelerating ICME design loops across academia and industry worldwide.

Read the original paper