Skip to content
AI.info

Research

HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

Overview Research area: Computer-use agents (CUAs) — agents that operate digital environments — with a focus on agent training, multimodal interaction, and reinforcement learning for tool use. Technic

HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents
arXiv
2609.38008
Published
2026-09-29
Authors
Tongbo Chen, Junbo Niu, Zhengxi Lu, Niu Lian, Fei Tang, Yuchen Yan, Yike Hong, Yong Du, Yizhou Liu, Bofan Chen, Yongliang Shen

AI summary

Overview

Research area: Computer-use agents (CUAs) — agents that operate digital environments — with a focus on agent training, multimodal interaction, and reinforcement learning for tool use.

Technical level: Advanced. The work involves a synthetic data-construction pipeline, supervised fine-tuning, and reinforcement learning with custom reward design, evaluated on OS-level agent benchmarks.

One-sentence scope: The paper argues that computer-use agents should interleave graphical user interface (GUI) actions with command line interface (CLI) commands, and presents a dataset, a two-stage training framework, and benchmark results for an agent built on that idea.

What This Paper Is About

Current computer-use agents take one of two approaches: they either act purely through the GUI, which is slow and mistake-prone, or they supplement GUI actions with application-specific APIs and tools, which require heavy engineering and do not scale across many applications. The authors propose combining GUI and CLI instead, since the GUI is general across software while shell commands are efficient. The central obstacle they identify is that existing models do not know when to reach for the CLI or how to use it correctly during a task, so the paper's goal is to teach that behavior through data and training.

Key Contributions

  1. A framing of the hybrid GUI-CLI paradigm. The paper argues that the next generation of computer-use agents should combine GUI interactions with command line interactions, rather than relying on the GUI alone or on application-specific APIs and tools.
  2. A data construction pipeline producing three trajectory types. The pipeline generates GUI-only, CLI-only, and interleaved GUI-and-CLI trajectories, resulting in the HybridCUA-8K dataset, which contains 5K hybrid trajectories and 3K verified RLVR tasks.
  3. A two-stage training framework. The method first applies supervised fine-tuning on the constructed trajectories, then reinforcement learning using the authors' CLI-aware rewards that push agents toward selective and reliable CLI use.
  4. An agent and cross-platform evaluation. HybridCUA-9B is evaluated on OSWorld and WindowsAgentArena, with reported gains over the base model on both.

Main Findings

  • Large gain on OSWorld: HybridCUA-9B reaches 53.6% accuracy on OSWorld, which the authors report as a 14.8 percentage point improvement over the base model.
  • Improvement on a second platform: The same approach improves performance on WindowsAgentArena by 4.0 percentage points, which the paper cites as evidence of cross-platform generalizability.
  • CLI-aware rewards shape behavior: The reinforcement learning stage uses rewards designed to encourage agents to use the CLI selectively and reliably, rather than indiscriminately.
  • Hybrid trajectories are constructible at scale: The pipeline yields 5K hybrid trajectories plus 3K verified RLVR tasks, suggesting hybrid GUI-CLI behavior can be generated without the per-application engineering that API-based augmentation requires.

Note on scope: only the abstract was available, so no further numbers, baseline comparisons, ablations, or dataset construction details are reported here. The abstract does not name the base model, the specific benchmark configurations, or how HybridCUA-9B compares against other published CUAs.

Methodology in Plain English

The authors start from the observation that agents need examples of good hybrid behavior before they can learn it, so they build a pipeline that generates training trajectories in three flavors: tasks solved purely through the GUI, tasks solved purely through the CLI, and tasks that switch back and forth between the two. This produces a dataset of 5K hybrid trajectories alongside 3K tasks that come with verifiable answers, suitable for reinforcement learning with verifiable rewards.

Training then happens in two stages. First, supervised fine-tuning exposes the model to those trajectories so it learns the basic pattern of combining GUI and CLI actions. Second, reinforcement learning refines the behavior using rewards the authors designed specifically around CLI usage, aiming to make the agent use shell commands when they help and avoid them when they do not. The resulting model, HybridCUA-9B, is then tested on OSWorld and WindowsAgentArena.

The abstract does not describe how trajectories were collected or synthesized, how the CLI-aware rewards are computed, or the specifics of the reinforcement learning setup.

Why This Matters

Impact on research: The paper challenges the assumption that scaling computer-use agents requires either GUI-only interaction or a growing library of application-specific tools. It offers an alternative direction — teach agents a general-purpose channel (the shell) and let them decide when to use it — and situates that decision-making as a learning problem addressed through data construction and reward design.

Real-world applications:

  • Operating-system and desktop automation: Agents that can configure settings, manage files, and run system utilities by mixing clicks with shell commands.
  • Software development and IT operations: Tasks where a GUI is used for inspection while build, test, and deployment commands run in a terminal.
  • Data and batch processing: Workflows where a CLI is far more efficient than repeated GUI actions, such as bulk file conversion or log analysis.
  • Enterprise support across heterogeneous software: Environments with many applications where writing and maintaining per-app integrations is impractical.

Industry relevance: Avoiding application-specific API integrations lowers the engineering cost of deploying agents across a broad software landscape. This is directly relevant to OS vendors, automation and RPA providers, and enterprise software teams that need agents to work across tools they do not control — and it introduces a safety-relevant question, since agents with shell access can take more consequential actions than agents limited to clicking.

Future Directions

  • Broader platform and application coverage: The paper demonstrates gains on OSWorld and WindowsAgentArena; whether the hybrid paradigm transfers to other operating systems, mobile interfaces, or specialized domains remains open.
  • Safety and control of shell access: Giving agents a command line raises questions the abstract does not address about guardrails, permissions, and preventing destructive commands — relevant given the emphasis on using the CLI "reliably."
  • Refining when to choose one modality over the other: The CLI-aware rewards are the mechanism for selectivity, but the abstract does not report how well-calibrated that decision-making becomes or what failure modes remain.
  • Scaling and generalizing the data pipeline: Whether the construction pipeline can produce larger and more diverse hybrid datasets, and whether it complements rather than replaces API- or tool-based augmentation, is left unanswered.

Target Audience

Researchers and engineers working on computer-use, GUI, and web agents; practitioners applying reinforcement learning to agent behavior and tool use; and product or platform teams evaluating whether agent automation should be built on GUI control, application APIs, or general shell access. Readers looking for implementation details of the reward design or the trajectory generation pipeline will need the full paper, since the abstract does not contain them.

Authors’ abstract

Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specific APIs or tools, which require substantial engineering effort and are difficult to scale across applications. We argue that the next generation of CUAs should combine GUI interactions with the command line interface (CLI), leveraging the generality of the GUI and the efficiency of shell commands. A critical challenge, however, is that current models do not know when or how to use the CLI during task execution. To address this challenge, we develop a data construction pipeline that produces three types of trajectories: GUI only, CLI only, and interleaved GUI and CLI trajectories. This pipeline results in HybridCUA-8K, containing 5K hybrid trajectories and 3K verified RLVR tasks. Building on these data, we propose a training framework with two stages: supervised fine tuning on the constructed trajectories, followed by reinforcement learning with our CLI aware rewards that encourages agents to use the CLI selectively and reliably. Experiments show that HybridCUA-9B achieves 53.6% accuracy on OSWorld, improving over the base model by 14.8 percentage points, and improves performance on WindowsAgentArena by 4.0 percentage points. These results demonstrate the effectiveness and cross platform generalizability of the hybrid GUI and CLI paradigm for computer use agents.

Read the original paper