Skip to content
AI.info

Research

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

Overview Research area: Machine learning / computer-use (GUI) agents — automated optimization of the runtime "harness" that surrounds a frozen GUI model. Technical level: Advanced. Scope: The paper in

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
arXiv
2610.00948
Published
2026-10-01
Authors
Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai

AI summary

Overview

  • Research area: Machine learning / computer-use (GUI) agents — automated optimization of the runtime "harness" that surrounds a frozen GUI model.
  • Technical level: Advanced.
  • Scope: The paper introduces GUI-HARVEST, an automatic harness optimizer that converts repeated multimodal GUI execution evidence into validated source-code edits, evaluated on OSWorld-Verified across six frozen backbones plus transfer to WindowsAgentArena.

What This Paper Is About

A GUI agent is not just a model — it is the model plus an executable "harness" that decides how screenshots and context are assembled, how clicks and keystrokes are executed, and how success is verified, recovered from, and terminated. The paper asks whether that harness can be improved automatically, while the model weights stay frozen, so the agent effectively gets better at computer use without retraining. The difficulty is that GUI failures are visual and noisy: a text trace may claim success while the screen shows an unresolved dialog, and the same task can succeed or fail across repeated runs.

Key Contributions

  1. Evidence-driven self-improvement for GUI agents. The authors formulate harness adaptation as evidence-driven improvement of an executable runtime around a frozen GUI model, with the model weights, benchmark tasks, evaluators, environment infrastructure, and resource limits all held fixed.

  2. Multimodal diagnosis and intervention validation. GUI-HARVEST converts repeated multimodal execution evidence into reusable code changes, organized around four optimizer roles: the Evidence Analyst (task-local diagnosis), the Cross-task Clusterer (recurring behavior modes), the Harness Engineer (source-aware edits with written predictions), and the Validator (score and behavior checks).

  3. Evaluation across backbones and environments. The paper demonstrates held-out Test gains across six backbones, performance at larger step budgets (15, 50, and 100 steps), and frozen-harness transfer beyond OSWorld to WindowsAgentArena, where the harness optimized on OSWorld is used without any WAA optimization.

  4. A validated promotion rule. Candidates are promoted only when they pass two deterministic code checks (L0 and L1), a utility hard gate (no decline on Search or Validation, strict gain on at least one), and a behavioral soft gate that checks whether the predicted behavioral change actually occurred.

Main Findings

  • Gains across all six backbones. At 15 steps under a unified protocol, Test scores improve by 1.49–10.16 percentage points across all six frozen backbones (Qwen3-VL-8B/32B-Instruct, OpenCUA-32B/72B, Gemini 3.1 Pro, GPT-5).

  • Largest single result. Qwen3-VL-32B-Instruct reaches 53.18% on Test (+10.16 points) and 50.94% on the full 361-task suite (+12.33 points), the largest full-suite gain reported.

  • Per-backbone full-suite results. Qwen3-VL-8B: 29.49 → 36.83 (+7.34). OpenCUA-32B: 30.71 → 34.24 (+3.53). OpenCUA-72B: 37.65 → 42.33 (+4.68). Gemini 3.1 Pro: 66.46 → 73.37 (+6.91). GPT-5: 55.20 → 62.42 (+7.22).

  • Execution variability is real and measured. In the baseline runs, 11.8–20.4% of Full tasks produced both zero and positive scores across three executions, depending on the backbone. This motivates treating repeated runs of the same task as a joint evidence unit.

  • Harnesses optimized at 15 steps keep improving at larger budgets. Without further optimization, Gemini 3.1 Pro reaches 79.14% at 100 steps. At 50 steps, the selected Qwen3-VL-32B-Instruct harness reaches 51.52%, versus a Qwen baseline of 32.60% (18.92 points).

  • Beats general harness optimizers on the same setup. From the same initial Qwen3-VL-32B-Instruct / Agent S3 harness, GUI-HARVEST gains 10.16 Test points, versus 2.00 for Self-Harness and 4.38 for Meta-Harness.

  • Beats a GUI-specific failure-driven optimizer's released harnesses. On OpenCUA-32B, GUI-HARVEST reports 34.31 Test / 34.24 Full versus LFF 31.83 / 32.12; on OpenCUA-72B, 44.18 / 42.33 versus LFF 41.29 / 39.50.

  • Transfers across platforms. Frozen harnesses transferred to WindowsAgentArena at 50 steps raise Qwen3-VL-32B-Instruct from 38.21% to 44.68% (+6.47) and GPT-5 from 50.88% to 64.75% (+13.87), with no WAA optimization.

  • Matched-backbone comparisons. The largest margins reported include 22.61 percentage points over CoAct-1 with GPT-5 (62.42% versus 39.81%) and 21.68 points over VLAA-GUI with Gemini 3.1 Pro (73.37% versus 51.69%) at 15 steps.

  • Accuracy–cost tradeoffs. With GPT-5, the 15-step harness nearly matches Agent S3's accuracy (62.4% versus 62.6%) at approximately 72% lower full-suite API cost ($72 versus $260). With Gemini 3.1 Pro, the 100-step harness reaches 79.1% at $125, versus Claude Sonnet 4.5 at 58.1% and $316. Qwen3-VL-32B-Instruct offers a lower-cost point of 50.9% at $34. The authors note these are source-reported operating points, not a shared-budget experiment.

  • Every component matters. Ablations on Qwen3-VL-32B-Instruct: removing visual evidence drops to 44.54% Test / 40.69% Full; removing the Cross-task Clusterer gives 45.66% / 42.13%; single-run evidence (K=1) gives 47.44% / 43.12%; hard-gate-only gives 49.36% / 48.08%; the full method gives 53.18% / 50.94%.

  • Behavioral checks change selection, not just scores. The hard-gate-only variant evaluates nine rounds yet selects a lower-scoring harness than full GUI-HARVEST, which terminates in round eight.

  • More repeats help only up to a point. In the sweep, K=3 gives the strongest final harness (50.9%), while K=5 increases rollout cost without improving the result (49.7%).

  • Useful interventions are backbone dependent. Qwen3-VL combines code routing with action recovery and completion checks; frontier models mainly benefit from routing, persistence, feasibility, and completion-policy changes; OpenCUA adaptations stay close to its native coordinate-action contract, emphasizing action normalization, file finalization, and terminal semantics.

  • Remaining failures are hard to see. Early edits address failures with explicit runtime signals such as repeated actions, invalid calls, or unsaved files. What remains increasingly involves successfully executed actions targeting the wrong object, route, or evaluator-relevant state — actions that may look valid in screenshots and logs.

Methodology in Plain English

The authors set up a loop in which a frozen GUI model runs each task three times (K=3) from clean environment snapshots, at a 15-step action limit, producing text, planned and executed actions, screenshots, tool outputs, scores, and termination metadata. Each group of runs is a "task-round bundle," accompanied by a Task Card (instruction, domain, applications, feasibility, hints) and a Harness Card (runtime and observable outputs). A bundle enters diagnosis when at least one valid run scores zero, and the other runs serve as comparison evidence.

Four roles then operate in sequence. The Evidence Analyst examines each bundle independently, comparing the terminal application state with the task objective and checking whether intended actions, executed actions, and subsequent screens agree; a deterministic verifier checks the harness version, cited run and step, quotation locality, and mechanically testable trajectory facts. The Cross-task Clusterer groups verified findings into "behavior modes" — recurring patterns shared by at least two tasks — after first assigning each finding a deterministic outcome category. The Harness Engineer then inspects and edits source code within a capability manifest, applies a bounded patch, runs local tests, and — critically — records a written prediction of the edit's observable behavioral effects before any evaluation. The Validator screens the candidate with L0 checks (edit permissions, diff limits, syntax, static rules) and L1 checks (imports, interface compatibility, unit tests, a smoke test), then runs it K times on search and validation tasks.

Promotion requires all four conditions: passing L0, passing L1, no score decline on Search or Validation with a strict gain on at least one, and a behavioral soft gate where each prediction needs at least one determinate verdict and more supporting than contradicting verdicts. Rejection rolls back to the current harness and its verified evidence. An update ledger records edits, predictions, scores, and decisions to guide later proposals. Optimization allows at most B=10 rounds with up to P=5 evaluated candidate edits per round, stopping when all five attempts at a harness state fail promotion. The test set stays sealed until optimization ends.

The evaluated backbones were Qwen3-VL-8B/32B-Instruct, OpenCUA-32B/72B, Gemini 3.1 Pro, and GPT-5, with Qwen3-VL and proprietary backbones starting from Agent S3 (default grounding model UI-TARS-1.5-7B, best-of-N disabled at N=1) and OpenCUA models using their official coordinate-action runtime. All four optimizer roles used Claude Sonnet 5 with the same role prompts across backbones. The benchmark was OSWorld-Verified with 361 tasks after excluding eight Google Drive tasks, split into 80 Search, 80 Validation, and 201 sealed Test tasks.

Why This Matters

Impact on research. The paper reframes self-improvement for GUI agents as a harness-search problem rather than a weight-training problem, and shows that generic harness optimizers built for math and software-engineering domains (Self-Harness, Meta-Harness) transfer poorly to GUI work. It also makes a methodological argument: because the same task can score both zero and positive across repeated runs (11.8–20.4% of Full tasks), single-trajectory diagnosis is an incomplete account, and repeated executions should be treated as the unit of evidence. The finding that interventions are backbone dependent — no universal patch — is a caution for anyone expecting one fix to generalize.

Real-world applications.

  • Desktop and browser automation for enterprise workflows, where an improved runtime can raise task completion rates without swapping or retraining the underlying model.
  • Cost-sensitive deployment of computer-use agents: the reported operating points suggest harness quality can reduce API spend (for example, $72 versus $260 for near-matched accuracy with GPT-5).
  • Software QA and regression testing, since harnesses can encode completion checks, persistence, and recovery behavior that catch unresolved states such as an open save dialog.
  • Accessibility tooling and robotic process automation built on GUI interaction, where verification and recovery logic matter as much as raw model capability.

Industry relevance. The practical bottleneck for computer-use agents is often reliability and inference cost, not raw model quality. This work shows that a frozen model plus a better runtime can deliver meaningful held-out gains and lower cost per task, which is directly relevant to teams shipping GUI agents on top of vendor APIs they cannot fine-tune. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.

Future Directions

  • Failures that leave no local signal. The paper notes that remaining errors increasingly involve successfully executed actions targeting the wrong object, route, or evaluator-relevant state, which can look valid in screenshots and execution logs. Detecting and correcting these without a clear runtime signal is an open problem.

  • Backbone-specific versus universal interventions. Since useful edits depend on the backbone — Qwen3-VL, frontier models, and OpenCUA each favor different changes — an open question is whether intervention libraries can be learned once and adapted, or whether each backbone must be optimized separately.

  • Beyond the two evaluated platforms. Transfer to WindowsAgentArena used only platform adapters changed. Whether the same evidence-driven loop generalizes to mobile, browser, or other GUI environments is not established.

  • Efficiency of the optimization loop. K=5 increased rollout cost without improving the final harness over K=3 in the reported sweep. Reducing the number of rollouts, rounds, or candidate attempts needed per promoted edit would make the approach cheaper to apply.

Target Audience

Researchers working on GUI agents, computer-use models, and self-improving agent systems will find the core contribution most relevant, as will practitioners who deploy agents on top of frozen commercial or open models and want better runtime behavior without fine-tuning. Readers interested in agent evaluation methodology will also benefit from the paper's treatment of execution variability and its two-gate promotion rule, and the accuracy–cost analysis speaks to teams making deployment decisions where inference budget matters. Some familiarity with harness/scaffold concepts and benchmark evaluation practice is assumed.

Authors’ abstract

The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.

Read the original paper