Research
QCoder Benchmark: Bridging Language Generation and Quantum Hardware through Simulator-Based Feedback
Overview Research area: Natural language processing, specifically LLM-based code generation and benchmark design, applied to quantum programming (Qiskit circuits targeting quantum hardware). Technical

- arXiv
- 2510.26101
- Published
- 2025-10-30
- Authors
- Taku Mikuriya, Tatsuya Ishigaki, Masayuki Kawarada, Shunya Minami, Tadashi Kadowaki, Yohichi Suzuki, Soshun Naito, Shunya Takata, Takumi Kato, Tamotsu Basseda, Reo Yamada, Hiroya Takamura
AI summary
Overview
- Research area: Natural language processing, specifically LLM-based code generation and benchmark design, applied to quantum programming (Qiskit circuits targeting quantum hardware).
- Technical level: Intermediate. The paper is readable without a quantum computing background, but the evaluation stages (transpilation, unsupported gates, circuit depth, state-vector fidelity) assume some familiarity with LLM evaluation pipelines.
- Scope: The paper introduces QCoder Benchmark, a dataset of 58 quantum programming problems with roughly 30 human-written submissions each and an average of 20 intermediate revision versions per code (1,740 problem–solution pairs total), together with a simulator-based evaluator that returns hardware-aware feedback for iterative code refinement.
What This Paper Is About
Existing code generation benchmarks are evaluated in software-only environments, where a program fails only if it throws a Python error or violates syntax. Quantum programming is different: generated code must also produce a circuit that respects hardware constraints such as which gates are supported and how deep the circuit may be. The paper addresses this gap by building a benchmark and evaluation service where quantum code is checked and executed against a simulated quantum device, then feeding the resulting constraint violations back to the model so it can revise its code.
Key Contributions
- A dataset of quantum programming contest problems paired with human-written solutions, collected from the QCoder platform, including revision histories that capture how human coders iterated toward working solutions.
- A simulator-based evaluator that goes beyond Python execution, checking in order of severity: Python runtime errors, use of unsupported quantum gates, circuit depth violations, and fidelity of the output state vector against a reference state.
- Iterative, feedback-driven code generators that convert the evaluator's structured report (for example
{"runtime_error": false, "gate_violation": false, "depth_violation": false, "state_match": true}) into a natural-language refinement prompt. - Empirical evidence that hardware-aware feedback improves LLM performance, and a comparison of LLM outputs against averaged human contest submissions.
Main Findings
- Reasoning models lead by a wide margin: Under baseline prompting with no refinement, o3 reached 65.52% success, compared with 18.97% for GPT-4o-mini, 10.34% for GPT-3.5-turbo, 29.31% for DeepSeek-R1-Distill-Llama-70B, and 10.34% for Qwen-1.5-14B-Chat.
- LLMs versus human coders: The averaged success rate of human-written contest submissions was 39.98%, so o3 (65.52%) outperformed the human average, while GPT-4o-mini and GPT-3.5-turbo fell below it.
- Failure modes differ by model: GPT-4o-mini's largest failure category was wrong output at 53.45%, followed by runtime errors at 17.24% and depth violations at 5.17%. GPT-3.5-turbo showed 36.21% runtime errors, 12.06% depth violations, and 31.03% wrong output. o3 showed 0.00% runtime errors, 1.72% depth violations, and 18.97% wrong output. Averaged human submissions showed 27.40% runtime errors, 10.03% depth violations, and 24.69% wrong output.
- Humans and weaker models share basic failure modes: Runtime errors were common for both GPT-3.5-turbo (36.21%) and human coders (27.40%), and depth violations were similar for human coders (10.03%) and GPT-3.5-turbo (12.06%).
- Refinement helps most early: Success rates improved across GPT-3.5, GPT-4o, and o3 with iterative refinement, with the largest gain from the first to the second iteration; gains became incremental after that.
- Improvement is not monotonic: Comparing o3 and human submissions across 1 to 15 refinement iterations, o3 showed occasional sudden gains (for example between iteration 4 and 5) but also stagnation or slight drops (for example from iteration 3 to 4), while human performance trended steadily upward with more refinements.
- Reported ceiling differs by section: The abstract states that reasoning-based models such as o3 reach up to 78% accuracy, while the baseline results table reports o3 at 65.52%; the paper does not state which iteration count corresponds to the 78% figure.
- Case study: For problem QPC001-A4, DeepSeek-R1's first attempt used
.initialize(), which is not in the allowed gate set and triggered a gate constraint violation; the second attempt passed runtime and constraint checks but produced the wrong output state; the third attempt succeeded using a different approach (rotations and controlled gates) from the reference solution. Many failure cases involved trivial issues such as missing imports (for exampleimport math).
Methodology in Plain English
The task is framed as: given a natural-language description of a quantum state to prepare plus constraints (allowed gates and maximum circuit depth), generate the body of a Python solve() function using Qiskit.
All models were compared using prompt-based generation, not fine-tuning, because the authors argue that prompt-based techniques matter most in domains where users are not NLP experts. Every model received the same prompt structure: a problem description, explicit constraints, and a code template with a solve() placeholder. Models were told to output only the function body, with no extra imports. Each model used its own default tokenizer and decoding settings, with no maximum token length.
Generated code was passed to the evaluator, which ran four checks in order of severity. First, Python execution to catch syntax and runtime errors, which halts the remaining checks. Second, transpilation and inspection for gates not supported by the hardware. Third, measurement of circuit depth against the problem's threshold. Fourth, execution on a simulator (Qiskit Aer in these experiments, replaceable by a real quantum computer) and comparison of the resulting state against the reference.
For the refinement condition, the evaluator's report was converted into a natural-language prompt. Labeled feedback covered wrong answers (WA), depth limit exceeded (DLE), unauthorized modules (UME), unauthorized quantum gates (UGE), and runtime errors (RE) with the error text included. Models revised their code for up to a fixed number of rounds (for example 3), or until all checks passed. Fine-grained failure rates were computed as the proportion of generations failing at each evaluation stage, and the same metrics were computed for the human submissions for comparison.
Why This Matters
- Research impact: The benchmark tests a class of generation where correctness depends on external hardware constraints, not just Python validity, and shows that simulator-derived feedback is an effective refinement signal. It also provides a human baseline (39.98%) that most non-reasoning models fail to beat.
- Robotics: Programs that control physical robots must respect timing and actuation limits, not merely run without exceptions.
- Embedded systems: Code for constrained devices must satisfy memory, resource, and device-compatibility limits that a standard interpreter cannot detect.
- Quantum computing toolchains: Developers writing Qiskit circuits need early warnings about unsupported gates and excessive circuit depth, which are the exact signals this evaluator surfaces.
- Education and competitive programming: Platforms hosting quantum programming contests can use the evaluator and revision histories to give structured feedback to learners.
Industry relevance: The work is directly relevant to quantum hardware and cloud providers whose SDKs impose gate sets and depth limits, to tooling vendors building automated code assistants for constrained hardware, and to teams evaluating whether to rely on reasoning-oriented models (o3 at 65.52%) rather than mid-tier models (GPT-4o-mini at 18.97%) for hardware-aware code generation. The paper states that the dataset and evaluation API will be released under a license prohibiting commercial use.
Future Directions
- Whether the feedback-driven framework generalizes to other domains with strict execution constraints, such as robotics and embedded system programming, which the authors explicitly leave to future work.
- Using a real quantum computer instead of the simulator for more precise feedback, an option the paper notes is available but does not evaluate.
- Explaining and stabilizing the non-monotonic refinement behavior of o3, which improved sharply at some iteration steps and stagnated or dropped at others.
- Extending beyond prompt-based generation to fine-tuning approaches, which the paper deliberately excluded from its comparison.
Target Audience
Researchers and practitioners in LLM code generation, benchmark design, and quantum computing tooling; developers evaluating LLMs for hardware-constrained programming; competition and education platform operators who want automated, hardware-aware feedback; and readers interested in how iterative refinement with domain-specific feedback differs from generic Python execution feedback.
Authors’ abstract
Large language models (LLMs) have increasingly been applied to automatic programming code generation. This task can be viewed as a language generation task that bridges natural language, human knowledge, and programming logic. However, it remains underexplored in domains that require interaction with hardware devices, such as quantum programming, where human coders write Python code that is executed on a quantum computer. To address this gap, we introduce QCoder Benchmark, an evaluation framework that assesses LLMs on quantum programming with feedback from simulated hardware devices. Our benchmark offers two key features. First, it supports evaluation using a quantum simulator environment beyond conventional Python execution, allowing feedback of domain-specific metrics such as circuit depth, execution time, and error classification, which can be used to guide better generation. Second, it incorporates human-written code submissions collected from real programming contests, enabling both quantitative comparisons and qualitative analyses of LLM outputs against human-written codes. Our experiments reveal that even advanced models like GPT-4o achieve only around 18.97% accuracy, highlighting the difficulty of the benchmark. In contrast, reasoning-based models such as o3 reach up to 78% accuracy, outperforming averaged success rates of human-written codes (39.98%). We release the QCoder Benchmark dataset and public evaluation API to support further research. (Codes and datasets are available at https://qcoder-bench.github.io/ )