Skip to content
AI.info

Research

RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

Overview Research area: Robotics benchmark design and evaluation of LLM-based coding agents on physical, physics-grounded engineering tasks. Technical level: Advanced (assumes familiarity with reinfor

RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers
arXiv
2609.34210
Published
2026-09-28
Authors
Haitong Ma, Chenxiao Gao, Rushi Qiang, Bo Dai, Na Li

AI summary

Overview

  • Research area: Robotics benchmark design and evaluation of LLM-based coding agents on physical, physics-grounded engineering tasks.
  • Technical level: Advanced (assumes familiarity with reinforcement learning, robot simulators, vision-language-action models, and agentic coding harnesses).
  • Scope: The paper introduces RLE-Bench, a 51-task benchmark organized into 4 workflows and 9 task families, and reports results for 11 model–harness combinations scored by an aggregate "RLE Index."

What This Paper Is About

Coding agents can already handle digital engineering work such as patching repositories and running machine-learning experiments, but it is unclear whether they can do the messier, physically grounded work of a robot-learning engineer: controlling real-time systems, training policies, reasoning from noisy sensors, and designing hardware. Existing robotics benchmarks mostly score a single artifact, such as a policy or a controller, rather than the broader set of design, debugging, and integration skills an engineer needs. The paper asks whether general-purpose coding agents can "qualify" for that job, and builds a benchmark to measure it.

Key Contributions

  1. A new benchmark spanning the robotics development stack. RLE-Bench contains 51 tasks in 9 families across 4 workflows: Interactive Control (Families 01–03), Policy Learning/Development (Families 04–05), Perception and Estimation (Families 06–07), and Mechanical Design (Families 08–09). Each task is defined as a tuple of instruction, development environment, budget, artifact contract, evaluation environment, and verifier.
  2. A standardized, hidden, executable evaluation protocol. All tasks ship as Harbor environments. Agents work in a resource-bounded development phase, then the submitted artifact crosses into a verifier image where hidden seeds, privileged states, and reference assets stay verifier-private; the verifier runs physical rollouts and produces the only authoritative reward report.
  3. An aggregate capability metric plus per-workflow profiles. Task scores in [0,1] are averaged across subtasks, then across families within a workflow, then across the four workflows to produce the RLE Index, in which each workflow family has equal weight. Leaderboards are accompanied by workflow-specific capability profiles and cost comparisons.
  4. Behavioral case studies across representative tasks. The paper analyzes scaffolding levels, development-time feedback, depth-sensor use under occlusion, and failure modes in mechanical design, highlighting capabilities and limitations that a single scalar score would hide.

Main Findings

  • Flagship models lead, with visual grounding as a differentiator. GPT-6 Astra and Claude Fable 5.1 outperform the other open- and closed-source models by a substantial margin, with the largest advantage in Interactive Control and Perception and Estimation — both requiring multimodal observation and repeated environment feedback.
  • Policy Development is the most evenly matched workflow. The performance gap between models is substantially smaller in Policy Development, which more closely resembles traditional machine learning engineering.
  • Physical consequences remain hard for everyone. In Mechanical Design, gaps are much smaller and no model consistently makes sound design decisions, especially where implementation choices have delayed or indirect physical consequences.
  • Scaffolding helps weaker models more than stronger ones. In Family01, three scaffolding levels were tested: L1 (basic robot and sensor APIs), L2 (adds SAM3 and Contact-GraspNet for perception and grasp planning), and L3 (adds privileged object and fixture positions). Richer scaffolding generally improved scores and reduced cost for most models, but GPT-6 Astra performed strongly at L1 and gained little or occasionally lost performance with more support.
  • Development-time feedback produces real but bounded improvement. In Family05, agents could query an evaluation service during development. In 29 of 36 sessions, the submitted version outperformed the agent's first development-time evaluation, and requested rollout episodes correlated positively with evaluation score. On open-design subtasks, development success predicted the official score within 5 points, but on robustness subtasks the official score fell 19 points below development on LIBERO and 41 points below on RoboTwin; several agents built their own perturbation proxies, but none closed the gap.
  • Depth helps, but handling occlusion separates models. In Family06, most models improved perception scores with RGB-D over RGB-only. High-performing models used depth to judge binary occlusion masks rather than only recovering surface or 3D positions. GPT-6 Astra trained a small convolutional neural network as a fallback when depth was unreliable, making it the only agent that actually used the GPU in the method-agnostic task.
  • Visible objectives can be met while physical consequences are missed. In the Family08 mobile base design task, most models passed structural validity checks, yet all nine models received zero worst-arm credit for static stability margin, indicating overlooked instability under load and collisions during motion.

Methodology in Plain English

The authors built a suite of executable robotics challenges in which an AI coding agent is dropped into a simulator workspace, given a plain-text instruction and a resource budget, and asked to produce a concrete deliverable — a controller, a trained model checkpoint, estimator code, or an MJCF mechanical design plus its controller. During a development phase the agent can write code, run it, observe simulated physical feedback, and iterate; the work is metered against budgets such as wall-clock hours, charged interaction steps, GPUs, and CPU counts. When development ends, the agent's artifact is moved into a separate verifier environment where the evaluation scenes, dynamics, embodiments, seeds, or tasks are hidden, and the verifier runs the physical rollouts that generate the authoritative score. Scores are normalized to [0,1] and averaged up a three-level hierarchy (task to family to workflow), then averaged across the four workflows with equal weight to give the RLE Index. Eleven model–harness combinations were run with reasoning effort set to "high" and web-search tools banned by default; harnesses included Claude Code, Codex CLI, Antigravity CLI, and Grok Build.

Why This Matters

The paper reframes robotics evaluation from "how good is this policy" to "can an agent do the engineering work." That shift matters because real robotics progress depends on building, calibrating, diagnosing, and repairing heterogeneous systems, not just training a single model.

Real-world applications:

  • Industrial automation: Family07 models a pick-and-place setting where a manipulator with two RGB-D cameras, force-torque sensors, and a magnetic gripper must clear U-shaped metal components from a bin onto a conveyor belt as fast as possible.
  • Humanoid locomotion and whole-body control: Family04 uses Unitree G1 motion tracking from LAFAN1 clips (29 joints, 50 Hz), relevant to legged robots that must survive perturbed dynamics, latency, and pushes.
  • Mobile manipulation hardware: Family08 covers a common base for Franka Research 3, XArm7, and UR5e arms that must reach shelf reference points while balancing payload, resource use, and stability.
  • Low-cost teleoperation hardware: Family09 adds gravity compensation to GELLO lead-arms, evaluated under unseen physical instances, poses, and payloads.

Industry relevance: The benchmark targets the practical question of whether agentic coding tools can be trusted with robotics engineering workflows, and it explicitly tracks cost alongside performance, which matters for teams deciding how much compute and scaffolding to allocate.

Future Directions

  • Closing the robustness gap. The paper reports that robustness subtasks, whose test-time perturbations are never shown to the agent, still drop official scores far below development scores; building robust policies is described as a remaining difficult challenge.
  • Broadening workflow coverage. The authors state that the four workflows in this version capture only part of the work of robot-learning engineers and that real-world robotics requires a broader range of capabilities.
  • Reducing simulation-to-reality uncertainty. Simulation performance does not establish real-world reliability or safety, and the scale and diversity of simulated scenes are described as insufficient to assess behavior across all deployment conditions.
  • Guarding against training-data contamination. The planned public release introduces a risk that future evaluations could be compromised by contamination.

Target Audience

Researchers and engineers working on coding agents, robot learning, and benchmark design; robotics teams evaluating whether agentic tools can handle engineering workflows; and machine-learning practitioners interested in how agent capabilities transfer from digital software tasks to physics-grounded, multimodal problems.

Note: numeric RLE Index values for the leaderboard are presented in Figure 1 and Figure 3 and are not given as numbers in the provided text, and the Appendix task card for Family06 is truncated mid-rubric.

Authors’ abstract

Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineering capabilities. Real-world robotics extends beyond control: agents must build, integrate, diagnose, and improve heterogeneous artifacts under resource constraints and reason from multimodal feedback. To evaluate these broader capabilities, we introduce RLE-Bench, a benchmark of robot-learning tasks spanning four representative robotics development workflows: interactive control, policy learning, perception and estimation, and mechanical design. We use diverse task-specific metrics to evaluate the artifacts submitted by the coding agents, from the success rate the agents achieved to the policy agents trained, the harness agent built, and the mechanical structures the agent designed. We aggregate these metrics into an overall RLE Index and report workflow-specific capability profiles, enabling systematic comparison of coding agents' capabilities across multiple capability dimensions. Beyond performance ranks, we also conduct in-depth case studies examining agent behavior on representative tasks, highlighting both current capabilities and limitations, and pointing to the opportunities robotics tasks have to offer for future agent training.

Read the original paper