Skip to content
AI.info

Research

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

Overview Research area: AI agents for software engineering — specifically, the intersection of computer-use agents (CUAs) that operate graphical interfaces and coding agents that edit source code. Tec

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
arXiv
2609.32600
Published
2026-09-26
Authors
Prince Zizhuang Wang, Chenhao Liang, Zelong Xu, Aojie Yuan, Xiaolin Zhou, Haiyue Zhang, Yue Zhao, Xiyang Hu, Shuli Jiang

AI summary

Overview

Research area: AI agents for software engineering — specifically, the intersection of computer-use agents (CUAs) that operate graphical interfaces and coding agents that edit source code.

Technical level: Intermediate. The paper is readable without deep machine-learning background, but it assumes familiarity with agent benchmarks, reinforcement learning framing (POMDP, reward, policy), and standard software engineering tooling. The formal notation is confined mainly to the problem formulation section.

Scope in one sentence: The paper introduces CUA-SWE, a benchmark of 105 tasks across four software engineering domains that requires agents to edit code, run and operate the resulting software through screenshots and graphical actions, and submit changes that pass deterministic behavioral tests.

What This Paper Is About

Existing coding agents and computer-use agents are studied largely in isolation, even though real development requires running software, observing how it behaves visually, and using those observations to decide what to change next. The paper asks two questions: how agents use GUI feedback to diagnose, repair, and verify software, and whether agents can complete software engineering tasks when required specification or operational information is available only through the running application's visual interface. To answer these, the authors build CUA-SWE — an environment, a benchmark, and an evaluation pipeline — where agents must modify code and configuration, execute commands, interact with running software, and inspect visual feedback within the same task.

Key Contributions

  1. An environment for software development through coordinated coding and computer use. Agents work with editable projects, development commands, mouse and keyboard actions, and screenshot observations, and they choose when to inspect source, run commands, interact with the software, and revise their implementation. The agent's action space is defined as the union of coding actions, computer-use actions, and a finish action that submits the final codebase for verification.

  2. A benchmark with deterministic evaluation and verifiable outcomes. CUA-SWE contains 105 tasks spanning four domains — 36 Web, 29 Game, 20 DevOps, and 20 Mobile. Each task provides an instruction, an initial codebase, a running application, and a verifier V_u(c) that combines three binary checks: requested functionality (V_u^task), protected existing behavior (V_u^reg), and restrictions on permitted changes (V_u^perm).

  3. An evaluation of frontier models and coding agents under two access conditions. The paper compares a code-only condition (source editing, builds, available tests, custom scripts, and their outputs) against a hybrid CUA condition that additionally provides screenshots and graphical interaction with the running application. Both conditions use the same instructions and the same protected patch tests.

  4. A supervised fine-tuning study as an early learning-signal probe. The authors fine-tune Qwen3.8-27B on action prefixes from verified reference replays with a 27/9 train/test split and measure the change in native-verifier pass rate.

Main Findings

  • Hybrid access helps frontier models overall. The paper reports that every frontier model in the four-domain comparison achieves higher aggregate task success with hybrid access than with code-only access, with GPT-6-Astra gaining 48.6 percentage points. The gains are reported as concentrated on tasks whose requirements include application materials, while source-specified subsets have near-zero domain-average changes and substantial variation across models. Table 1 also contains individual cells where hybrid is lower than code-only — for example, Claude-Sonnet-5 on Web (8.3 code-only vs. 0.0 hybrid) and on Game (6.9 code-only vs. 3.4 hybrid) — and one em dash for Claude-Opus-5 on Mobile, indicating no reported result.

  • GPT-6-Astra leads the four-domain comparison. It records 59.9% mean task success, 17.7 percentage points above GPT-5.6-Sol, and achieves the highest observed hybrid success within every domain. Its reported per-domain hybrid scores are 66.7 (Web), 37.9 (Game), 80.0 (DevOps), and 55.0 (Mobile); its code-only scores are 27.8, 17.2, 0.0, and 0.0 respectively.

  • Leader advantage concentrates in precise recovery of relationships and rules. On the Mobile gear task, GPT-6-Astra reconstructs all 23 driving relations and implements their angle, phase, and editing behavior, while GPT-5.6-Sol selects the wrong driving rim at two compound assemblies. In sprite animation, GPT-6-Astra recovers the complete pose sequence whereas GPT-5.6-Terra's near-complete reconstruction retains incorrect pose assignments. In DevOps, GPT-6-Astra uniquely solves the gauge-calibration task by probing endpoint behavior and updating the inferred mode when the deployment changes.

  • Following the consequences of a change distinguishes agents. In the nested-selection task, GPT-6-Astra first repairs resizing, then drags the object beyond its containing frame; when selection is lost it extends hit testing to selectable descendants and repeats the interaction. GPT-5.6-Sol and Claude-Opus-5 repair the resize behavior but leave the selection defect in their submitted patches.

  • Broader coverage coexists with economical interaction on shared successes. GPT-6-Astra's mean episode time is 4.3 minutes versus 3.7 minutes for GPT-5.6-Sol. On Web tasks solved by both GPT-6-Astra and each peer (18 tasks per comparison), GPT-6-Astra uses fewer responses on 83.3% of shared successes with Grok-4.6 and 72.2% with Claude-Opus-5.

  • Retries broaden coverage; consistency across attempts measures reliability. In the three-attempt Game study, GPT-6-Astra solves 34.5% of tasks in all three attempts, compared with 3.4% for GPT-5.6-Sol. On Mobile, later attempts add a sprite repair for GPT-5.6-Sol and a schedule repair for GPT-6-Astra.

  • Supervised fine-tuning shows an early positive learning signal. Training on gold-patch replays raises the native-verifier pass rate from 11.1% to 33.3%, measured with one greedy attempt per task and matched interaction settings. This measure is patch correctness P_i; benchmark task success additionally checks trajectory validity U_i.

Methodology in Plain English

The authors start from concrete software behaviors that can be reproduced in a running application and whose implementation can be modified in the codebase. An LLM-assisted authoring process constructs the instruction, the initial faulty codebase, a reproducible runtime setup, a gold reference repair, plausible negative repairs, and protected behavioral tests. A task is admitted only after human review and executable validation of its requirement, runtime behavior, reference solution, and verifier. The verifier must at minimum return 0 for the original faulty program and 1 for the gold repair, and must reject negative repairs that satisfy only part of the requirement — for example, a repair that clears a masked date field in both the valid-clearing and invalid-input cases.

The benchmark also verifies the interactive development path: that the application can be driven to the intended state, that relevant visual observations are available, that code changes take effect during development, and that verification is reproducible from a clean project state. Development-time information and evaluation-time information are deliberately separated, so the agent reasons from its instruction and interaction history while the evaluator may use privileged application state.

Tasks are organized by where their requirements are specified: source-specified (S) tasks, where the instruction and readable project code define the target behavior, and application-material-dependent (M) tasks, which require interpreting additional materials such as imported drawings, service contracts, or graphical reference cards. Construction-time agent trials calibrate difficulty and generate controlled variants and task families, but correctness is always decided by the fixed verifier.

Evaluation runs each task under two conditions. Code-only agents may read and edit source, run permitted builds, available tests, and custom scripts, and inspect outputs; graphical actions, served application access via development servers, direct HTTP or sockets, and browser automation are excluded. Hybrid CUA additionally starts the application and lets the agent observe pixels and issue mouse, keyboard, and navigation actions. Direct DOM reads, accessibility trees, in-browser script evaluation, and direct HTTP extraction remain outside the observation interface. Task success is defined as S_i = P_i · U_i, where P_i is patch correctness and U_i checks that tool access and observation evidence match the assigned condition; the reported metric is TSR = 100 · |E|⁻¹ · Σ P_i·U_i, with the four-domain mean giving Web, Game, DevOps, and Mobile equal weights at one attempt per task.

Why This Matters

Impact on research. The paper argues that coding benchmarks and computer-use benchmarks assess complementary capabilities in isolation, and that connecting the two is necessary to study how developers actually work. It provides deterministic, executable correctness criteria for the resulting software plus a verifiable reward signal (R_u(s_T) = V_u(c_T) ∈ {0,1}) intended for post-training. It also frames the whole problem as a finite-horizon partially observable Markov decision process where a single vision-language policy selects both coding and computer-use actions, giving a formal handle for future agent training work.

Real-world applications:

  • Frontier agent evaluation for integrated development workflows — measuring whether an agent can repair a game's collision behavior, a web form's input semantics, a mobile route planner, or a DevOps service configuration, not just patch a repository.
  • Diagnosing runtime failures that static code inspection misses — such as the Captain Callisto task where one pickup incorrectly changes the inventory from 0/2 directly to 2/2, or the Campus Pocket task where a faulty route uses a closed H6 lift while the reference repair follows H3, H4, and H5 for a total of 74 m.
  • Recovering specification from artifacts rather than text — the Mobile gear task requires reconstructing connections from received assembly drawings before implementing the corresponding motion, mirroring situations where requirements arrive as diagrams, service contracts, or reference cards.
  • Generating supervision data for agent post-training — verified reference replays supply action prefixes for supervised fine-tuning, with the executable verifiers proposed as a reward source for RLVR.

Industry relevance. The evaluated set includes frontier CUA models and harness systems such as Codex and Claude Code, and the comparison of code-only versus hybrid CUA access speaks directly to how much value an agent vendor's GUI access adds for development tasks. The three-attempt results (pass@1 versus pass@3, where the difference measures the task coverage added by additional attempts) are relevant to reliability and retry budgeting, and the interaction-economy findings bear on inference cost per solved task.

Future Directions

  • Establishing a verifiable learning signal and scaling RLVR training in this environment. The authors list this as immediate future work, noting that it requires efficient parallel rollout collection, stable optimization over long multimodal trajectories, and effective credit assignment across code edits, observations, and interactions.
  • Broadening benchmark coverage. The current benchmark spans Web, Game, DevOps, and Mobile; expanding to more domains, applications, and tasks is proposed to widen coverage of visual software engineering.
  • Supporting richer, longer workflows. Future extensions are described as including development workflows where agents coordinate coding and computer use across multiple tools and applications over longer horizons.
  • Understanding capability boundaries more precisely. Open questions raised include why local corrections can leave connected behaviors unresolved (as in Core Ball event ordering and Mobile schedule dependency reconstruction), and how consistency across repeated attempts should be separated from additional task coverage.

Target Audience

The paper is most useful to researchers and engineers building or evaluating autonomous software agents — particularly those working on computer-use agents, coding agents, or hybrid systems that combine both. It is also relevant to benchmark designers interested in deterministic, executable verification rather than LLM-as-judge scoring, and to practitioners deciding whether to give an agent graphical access to a running application as part of its development loop. Readers studying agent post-training will find the SFT setup and the proposed RLVR extension directly applicable, while readers focused on software engineering practice may find the four-domain task taxonomy (Web, Game, DevOps, Mobile) and the source-specified versus application-material-dependent distinction useful as a way to classify what agents can and cannot currently do.

Authors’ abstract

Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and evaluation pipeline for software engineering with computer use. Beyond studying how GUI feedback supports diagnosis and repair, we ask whether agents can complete software engineering tasks when required specification or operational information is available only through the running application's visual interface. CUA-SWE spans four software engineering domains and requires agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback within the same task. Each task includes deterministic, task-specific tests that verify whether the resulting software satisfies the requirements and preserves specified behavior. Our evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes. We examine performance across domains and task information requirements, alongside the development behaviors associated with successful repairs. CUA-SWE provides a unified testbed for studying how agents use visual feedback and interaction to guide software engineering, with executable correctness criteria for the resulting software.

Read the original paper