Research
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
Overview Research area: Computer vision and GUI (graphical user interface) agent research, specifically synthetic training-data generation using pretrained image generators as visual world models. Tec

- arXiv
- 2610.01215
- Published
- 2026-10-01
- Authors
- Cheng Yang, Yifan Wu, Yutao Huang, Zhaohua Zhang, Beiduo Chen, Muxi Chen, Chenchen Zhao, Hexuan Deng, Haolin Yang, Geyuan Zhu, Sa Zhu, Jianhuan Zhuo, Qiuyong Xiao, Jianhao Ruan, Yiran Peng, Jiayi Zhang, Tian Ye, Xinlei Yu, Tianwen Jiang, Jihong Zhang, Yuyu Luo
AI summary
Overview
Research area: Computer vision and GUI (graphical user interface) agent research, specifically synthetic training-data generation using pretrained image generators as visual world models.
Technical level: Intermediate. The paper assumes familiarity with GUI agents, screenshot-based policies, supervised fine-tuning, and embedding-distance metrics, but its central idea (generate training screenshots instead of running the software) is stated plainly.
Scope: The paper describes AutoGUIWorld, a framework that synthesizes GUI interaction trajectories with a planner and an image generator instead of deploying real software, and reports how fine-tuning on those trajectories transfers to real desktop, professional-interface, and scientific-software benchmarks.
What This Paper Is About
Training GUI agents requires large numbers of interaction trajectories showing how software responds to clicks, typing, scrolling, and dragging, but collecting those trajectories is limited by which applications can be installed, configured, and run reproducibly. AutoGUIWorld asks whether a pretrained image generator, guided by a task planner, can stand in for the software environment itself and produce screenshots of the next GUI state after each action. The goal is to generate spatially annotated training trajectories without deploying or running the corresponding software, then test whether agents trained on that synthetic experience perform better on real desktop and scientific tasks.
Key Contributions
-
GUI trajectory generation without execution. AutoGUIWorld combines task planning, scene-conditioned task generation, and image generation to synthesize screenshot-action-screenshot trajectories without deploying or running the corresponding software environments.
-
A curated training corpus with coverage analysis. The framework produces 79,266 grounded and filtered step-level samples across Ubuntu, Windows, macOS, and Chrome, with an accompanying analysis of interface coverage, visual alignment in embedding space, and transition-level defects.
-
Transfer to real environments. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories (the resulting model is called AGW-35B) improves all four interactive benchmarks, including OSWorld from 33.0% to 40.8% mean task score and ScienceBoard from 14.0% to 32.2% task success.
-
A transition-level quality-control pipeline. The paper documents a VLM audit over 42,526 evaluable desktop transitions that flags 1,293 action–image inconsistencies (3.04%), with 818 of 39,351 transitions remaining inconsistent (2.08%) after removing steps with an invalid action, absent target, or incorrect grounding.
Main Findings
-
Task execution improves across all four interactive benchmarks. The equal-weight mean of the four benchmark scores rises from 23.6% to 36.5% (+12.8 points) from the base Qwen3.5-35B-A3B model to AGW-35B at checkpoint 469. macOSWorld and ScienceBoard show the largest gains, at 16.9 and 18.2 points.
-
Desktop application scores rise. OSWorld reaches 40.8% from 33.0%. System tasks gain 25.0 points and GIMP gains 19.2 points, while Calc and Impress improve by 8.5 and 12.8 points. Multi-app tasks rise from 12.0% to 20.9% and contribute the largest share of the overall OSWorld gain. On Windows Agent Arena, the score rises from 19.4% to 27.9%, with Writer and VLC gaining 26.3 and 23.8 points; six domains remain unchanged.
-
Two domains decline. Chrome on OSWorld falls from 39.0% to 26.0%, and Chrome on WAA falls from 11.2% to 0.0%.
-
macOSWorld success rises substantially. Success goes from 28.1% to 45.0%. System & Interface improves from 31.0% to 62.1%, Advanced from 13.3% to 40.0%, and System Apps from 44.7% to 65.8%. Multi-app success doubles from 13.8% to 27.6%. Safety decreases from 20.7% to 17.2%.
-
Scientific software operation improves. ScienceBoard success rises from 14.0% to 32.2%, with gains in all five software domains. KAlgebra improves from 6.5% to 48.4% and ChimeraX from 37.9% to 55.2%; together they account for 18 of the 26 net additional successful tasks, as the total increases from 20 to 46. Celestia, GRASS GIS, and TeXstudio gain 9.1, 11.8, and 6.3 points. Across the four benchmarks, 25 of 35 domains improve, seven remain unchanged, and three decline.
-
Visual grounding in professional interfaces improves. On ScreenSpot-Pro, overall accuracy rises from 31.7% to 57.1% (+25.4 points) across all six domains. Icon accuracy rises from 16.2% to 32.9% and text accuracy from 41.2% to 72.0%. Office shows the largest gains for both target types (24.5 points for icons, 41.2 points for text). Scientific text grounding improves by 40.3 points to 79.9%, and Creative text grounding by 33.8 points to 74.7%. Icon gains range from 12.4 points in Dev. to 24.5 points in Office.
-
Synthetic data exceeds the real-demonstration comparison here. AGW-35B uses approximately 80k synthetic examples, while the AgentNet run uses 350k examples and continues through step 686. AGW-35B at step 469 exceeds the best evaluated AgentNet checkpoint on all four benchmarks, with the largest gap on ScienceBoard (32.2% versus 16.8%).
-
Gains appear early and persist. AGW-35B improves on every benchmark by step 80 and reaches its highest score on all four at step 469. ScienceBoard increases at each evaluated checkpoint; macOSWorld continues gaining after step 157; OSWorld and WAA fluctuate between checkpoints. The four-benchmark mean increases at every evaluation, from 23.6% for the base model to 36.5% at the final checkpoint.
-
Visual feature alignment is on the scale of real cross-source variation. Qwen MMD normalized distance (rho) ranges from 0.59 to 1.24 across the four domains. The matched distance is below the real cross-source reference on macOS and Windows; Web and Ubuntu are 14% and 24% above the real cross-source median. ScaleCUA group-split floors are 0.109 on Ubuntu, 0.003 on Windows, 0.007 on Web, and 0.004 on macOS.
-
Source identity is still detectable. C2ST AUC is approximately 1.00 for synthetic–real pairs and 0.95–1.00 for real–real pairs. A leave-one-real-source test has mean AUC 0.996 for Ubuntu and 0.976 for Web, compared with real-source controls of 0.960 and 0.925; Windows and macOS reach 0.990 and 0.912.
-
Low-level residuals are domain-specific and small in content-matched regions. In the resolution-matched Ubuntu and Windows domains, text and icon C2ST AUC falls between 0.44 and 0.57, near chance. Ubuntu flat backgrounds retain an AUC of 0.65, the clearest remaining low-level residual.
-
Detected interface coverage is higher than the real reference. AutoGUIWorld exceeds ScaleCUA in all twelve domain–threshold comparisons, with relative gains of 63–77% on Ubuntu, 51–75% on Windows, 49–71% on Web, and 10–13% on macOS. On Windows at confidence 0.15, 142 synthetic elements cover 0.351 of the screen versus 154 elements covering 0.212 in ScaleCUA. A manual audit of 180 detection overlays found no systematic tendency to label generated texture as UI elements.
-
Trajectories vary in length and action mix. Ubuntu and Windows average 11.9 and 11.4 steps with P90 lengths of 24 and 25; macOS averages 7.1 steps and 2.50 open applications and has the highest action-type entropy at 0.92. Rollout failure is 0% on Ubuntu, 0.29% on Windows, and 0.68% on macOS. Precondition and blocker fields are zero throughout this release, and matched real-trajectory metadata are unavailable for this analysis.
Methodology in Plain English
The researchers build a pipeline with two separate layers. A planning layer (a "meta planner") takes a task instruction and a fixed seed context and writes out an ordered list of atomic GUI actions, such as clicking, typing, scrolling, or dragging, along with a description of the visual change each action should cause. A visual world-model layer (the image generator Image2) then turns those descriptions into actual screenshots.
The process works as follows. First, an initial GUI scene is sampled from a structured specification covering three things: the operating system and platform conventions, the visual appearance (theme, palette, wallpaper, typography), and the initial interface state (window count, layout, foreground application, visible controls). That specification is compiled into a detailed description and rendered into a seed screenshot.
Second, a task is generated conditioned on that seed, so the instruction only refers to applications and elements actually present on the screen. Third, the meta planner produces the action sequence and expected visual consequences. Fourth, a component called Voyager walks through the plan: it looks at the current clean screenshot and the planned action, produces a first-person thought, an action summary, and an "after-action" rendering prompt, and Image2 edits the current frame into the next screenshot. Each new screenshot becomes the reference for the next step, which is how layout and background content carry through a trajectory.
Pointing actions need spatial labels. A tool called LocateAnything finds the target element on the pre-action screenshot, and the center of that region becomes the action coordinate. Clean screenshots are kept separate from annotated frames so the policy input is never contaminated by markers.
Before training, a quality-control pass audits the transitions with a vision-language model to catch cases where the visible screen change does not match the action that was supposed to cause it. The final corpus is converted into single-step training instances: clean screenshot plus task instruction plus history as input, action plus point as target.
To evaluate, the team fine-tunes Qwen3.5-35B-A3B on the synthetic trajectories, tests six checkpoints from one run on five benchmarks, and reports mean task score on OSWorld and Windows Agent Arena (retaining partial credit) and task success rate on macOSWorld and ScienceBoard. ScreenSpot-Pro uses point-in-box accuracy, with unparseable predictions counted as errors.
Why This Matters
Impact on research. The paper tests the proposition that pretrained image-generation priors, combined with task planning and transition-level filtering, can substitute for running software during data collection. If the reported transfer holds, the bottleneck for GUI agent training shifts from "which applications can we install and script" to "which applications can we describe." The paper also supplies a measurement apparatus — Qwen embedding MMD, C2ST source discrimination, DINOv2 cross-checks, content-matched ROI diagnostics, and OmniParser coverage — for deciding how close synthetic screenshots are to real ones, and it reports where the residual gaps are (Ubuntu flat backgrounds, Ubuntu overall visual distance, text entry and dragging transitions).
Real-world applications:
-
Agent training for platforms with hard-to-deploy software. Scientific, CAD, and electronic-design tools are called out in the paper as exposing specialized states and operations that carry installation, configuration, and licensing costs; the ScienceBoard results are the direct test of this use case.
-
Coverage of interface states that are rare in collected data. Because the initial seed is sampled from a structured space of window counts, layouts, foreground relations, and visible controls, the method can deliberately produce layouts that are underrepresented in real corpora.
-
Data augmentation alongside real demonstrations. The AgentNet comparison suggests a role for synthetic trajectories as a complement or partial substitute when real trajectories are expensive, though the two runs use different inference settings.
-
Catching defective transitions before training. The VLM audit and quality-control pipeline is reusable on its own for anyone generating or curating screenshot trajectories, including from real recordings.
Industry relevance. The work speaks to teams building computer-use agents for desktop operating systems and professional software, where the cost of maintaining reproducible execution environments is a direct constraint on data scale. The reported per-domain results also give a practical map of where synthetic training helps most and where it hurts: multi-app tasks improve but remain below many single-application domains, and Chrome regresses on both OSWorld and WAA.
Future Directions
-
Reducing long-horizon error accumulation. The discussion states that as trajectories grow longer and interactions become more involved, generation errors may accumulate into hallucinated interface states or action outcomes. The reported 2.08% residual inconsistency rate after filtering is measured on desktop transitions, and the paper notes residual errors concentrate in text entry, dragging, scrolling, and keyboard interactions, including command substitution, invented document content, premature formatting, missing action effects, and unintended persistent-content changes.
-
Training domain-specific visual world models. The discussion proposes training these models on domain-specific interactions so they better capture the interface structures, operation rules, and state transitions of particular environments, yielding "vertical GUI world models" tailored to particular software or workflows for better long-horizon consistency.
-
Explaining and fixing the Chrome regressions. Chrome falls on OSWorld (39.0% to 26.0%) and WAA (11.2% to 0.0%), and MacOSWorld Safety falls from 20.7% to 17.2%. The paper reports these declines but does not attribute them to a cause.
-
Comparing against real trajectories under matched conditions. The synthetic-versus-real comparison uses different inference settings (Appendix C.1), and the paper also notes that matched real-trajectory metadata are unavailable for its trajectory-structure analysis. A controlled comparison would clarify how much of the gap and advantage is attributable to the data versus the setup.
-
Enriching trajectory annotations. Precondition and blocker fields are zero throughout the current release, and rendered terminal states are almost always successful, so the corpus currently provides little supervision for preconditions or failure recovery.
Target Audience
This paper is most useful to researchers and engineers working on GUI agents and computer-use models who need training data at a scale their execution infrastructure cannot supply, and to practitioners evaluating synthetic data as a substitute or complement for human demonstrations. It is also relevant to anyone studying world models for digital environments, since it treats an off-the-shelf image generator as an action-conditioned transition model and reports the fidelity limits of that substitution. Readers interested in evaluation methodology — synthetic-versus-real distribution comparison, source-discrimination probes, and automated transition auditing — will find the measurement suite as valuable as the headline benchmark numbers.
Authors’ abstract
GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.