Research
AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
Overview Research area: GUI agents / mobile virtual assistants, benchmark quality and data curation, reinforcement learning for vision-language models. Technical level: Intermediate. The paper combine

- arXiv
- 2510.18488
- Published
- 2025-10-21
- Authors
- Ho Fai Leung, Xiaoyan Xi, Fei Zuo
AI summary
Overview
- Research area: GUI agents / mobile virtual assistants, benchmark quality and data curation, reinforcement learning for vision-language models.
- Technical level: Intermediate. The paper combines benchmark auditing with reinforcement-learning training, and its core argument can be followed without deep mathematics, though the methodology sections include policy-gradient and reward-shaping formulations.
- Scope (one sentence): The paper argues that the widely cited ~60% success ceiling for GUI agents on AndroidControl is largely an artifact of benchmark flaws, and it supports that claim with a purified benchmark (AndroidControl-Curated) and a 3B-parameter model (Magma-R1) trained on 2,400 curated samples.
What This Paper Is About
On-device assistants such as Siri and Google Assistant are limited to simple tasks and depend on developers explicitly supporting structured APIs like Apple's App Intents or Google's App Actions. GUI agents, which read the screen and tap or type directly, could automate any app without developer support, but they are widely considered unviable because even the best models (e.g. Qwen3-VL-235B) appear capped at roughly 60% success on benchmarks like AndroidControl. This paper traces that ceiling not to model limitations but to the benchmark itself — roughly 30% of AndroidControl contains ambiguities, unaccounted-for valid actions, and factual errors — and rebuilds the benchmark to measure agents fairly.
Key Contributions
- Systematic flaw analysis: The authors identify and characterize critical deficiencies in a mainstream GUI agent benchmark (AndroidControl), arguing that poor benchmark quality — not model capability — is a key bottleneck in the perceived viability of on-device GUI agents.
- A reproducible purification pipeline: They propose a semi-automated, two-stage method combining bounding-box-based grounding evaluation with LLM-assisted review, correction, and human expert verification.
- Release of AndroidControl-Curated: A purified benchmark on which existing models score far higher than previously believed, including a reported 76.5% success rate for Qwen3-VL-235B (a 15% increase) on the Hard subset.
- Release of Magma-R1: An open-source model, based on a 3B-parameter multimodal model, post-trained on just 2,400 curated samples with 60 hours on an H20 GPU (approximately $60), reported to match models trained on 13X more data.
Main Findings
- The ~60% ceiling was a benchmark artifact: The paper reports that approximately 30% of AndroidControl suffered from ambiguities, multiple plausible solutions, and factual inaccuracies that systematically penalized correct model behavior.
- Grounding evaluation was too strict: The original benchmark used exact point matching with a tolerance threshold often set to 0. Replacing this with bounding-box intent alignment (whether the predicted point falls inside the bounding box of the target UI element) produced large gains on its own — for example, Qwen3-VL-235B's Hard success rate rose from 61.2% to 71.7% (+10.5) from this single change, and Magma-R1's from 57.6% to 69.1% (+11.5).
- Task-level correction added further gains: Moving from AndroidControl-Curated-Box to the fully corrected AndroidControl-Curated added +4.8 for Qwen3-VL-235B (to 76.5%), +6.2 for Magma-R1 (to 75.3%), +4.3 for GUI-R1-7B (to 57.5%), +3.1 for Infi-GUI-R1 (to 70.7%), and +5.0 for GUI-R1-3B (to 54.4%).
- Three deficiency categories dominate: Wrong Ground Truth accounted for 37.63% of samples, Unclear Task for 24.13%, and Multiple Valid Actions for 8.12%; the paper states these collectively affect over 70% of AndroidControl samples.
- Small models close the gap on the purified benchmark: On the Hard subset, Magma-R1 (3B) reaches 75.3% success rate versus 76.5% for Qwen3-VL-235B — a model the paper describes as 200 times larger in parameters.
- Magma-R1 leads the Easy subset: In Table 1, Magma-R1 records 88.0% success rate on AndroidControl-Curated-Easy (Type 91.3%, Grounding 94.2%), ahead of the next-best listed score of 80.6% for OS-Atlas-4B.
- Infi-GUI-R1 is competitive despite far more training data: Trained on over 31k original AndroidControl data points, it reaches 70.7% success rate on the Hard subset (Type 78.5%, Grounding 72.8%).
- Proprietary GPT-4o struggles on grounding: GPT-4o is listed at 19.4% success rate on Easy and 20.8% on Hard, with grounding accuracy of 0.0% on both, under the bounding-box evaluation.
- Type accuracy improved alongside success: Under the full purification, Qwen3-VL-235B's Type score on Hard moved from 67.3% to 88.2%, and its Grounding score from 78.3% to 83.6%.
- The paper reports one internal inconsistency: The introduction states Infi-GUI-R1-3B exceeds 70% with an 11% increase, while Table 1 lists Infi-GUI-R1 at 70.7% on the Hard subset. The number of samples removed or retained in the final curated benchmark is not reported.
Methodology in Plain English
The work has two halves: fixing the benchmark, and training a model on the fixed data.
Fixing the benchmark (two stages). First, the authors change how a correct click is judged. Instead of requiring the predicted coordinate to match the labeled coordinate almost exactly, they use the app's DOM or accessibility tree to find the bounding box of the UI element that contains the labeled point, then check whether the model's predicted point falls anywhere inside that box. This is done before any data is edited, producing an intermediate version called AndroidControl-Curated-Box. Second, they hunt for bad data. They run three strong agents — Qwen3-VL-235B, InfiGUI-R1, and GUI-R1 — over the benchmark and flag every task that all three fail, on the assumption that a task no capable agent can solve is more likely flawed than universally hard. Each flagged task, along with its context and the failed execution traces, is passed to a large language model acting as an automated reviewer, which assigns a deficiency category from a defined taxonomy, proposes a revised instruction, a revised ground-truth trajectory, and a written rationale. Human experts then verify a large random sample of these automated corrections before anything is applied. The result is AndroidControl-Curated, covering both the Easy and Hard subsets.
Training Magma-R1. The model starts from a 3B-parameter multimodal base and is trained with a policy-gradient reinforcement learning method the authors call Generative REINFORCE with Policy Optimization (GRPO), which constrains how far the policy can move per update, in the spirit of PPO. Advantages are computed by normalizing rewards within each batch. Two adjustments address known GUI problems. To avoid sparse rewards, grounding reward is a smooth 2D Gaussian kernel over the distance between the predicted box center and the ground-truth point, so nearly-correct clicks get partial credit instead of zero. To avoid the long-tailed dominance of click actions, training batches are built by stratified sampling over action types toward a more balanced target distribution rather than by purely random sampling. Training used 2,400 curated samples and 60 hours on an H20 GPU, at an approximate cost of $60.
Why This Matters
Impact on research. The paper makes an argument that applies well beyond GUI agents: if a benchmark systematically penalizes correct behavior, then measured progress, model comparisons, and training data selection are all distorted. The authors report that training on larger benchmark datasets sometimes reduced accuracy, which they attribute to the benchmark teaching incorrect behaviors. If their findings generalize, benchmark auditing becomes a first-class research activity rather than a preliminary chore.
Real-world applications:
- In-car assistants: The authors situate the work in automotive contexts, where hands-free navigation, music, and vehicle control matter for safety, and where the affiliation (BMW ArcherMind Information Technology Co. Ltd.) reflects that interest.
- On-device assistants on phones: Models in the 1.8–3.25B range (Google's Gemini Nano) and Apple's 3B parameter foundation model are the target deployment class, promising fast, private, offline interaction.
- Automating unsupported and legacy apps: Because GUI agents act on the interface rather than on published APIs, they can reach apps whose developers never built assistant integrations.
- Enterprise app automation: Reading a screen and issuing taps and keystrokes generalizes to form filling and multi-step workflows in internal tools that lack API coverage.
Industry relevance. The cost and size claims are the commercially salient part: a 3B model trained for approximately $60 on 2,400 samples is reported to perform comparably to a 235B model, suggesting on-device deployment is closer than the previous benchmark numbers implied. The authors release both the benchmark and the model to encourage adoption of the purified evaluation.
Future Directions
- Extend purification to other benchmarks. The pipeline is presented as reproducible and semi-automated, but this paper applies it only to AndroidControl; whether AndroidControl-Curated-Box-style grounding fixes and task-level corrections transfer to other GUI benchmarks is untested here.
- Scale the automated review beyond consensus failure. Flagging tasks that all three expert agents fail will miss flawed tasks that a strong agent happens to solve by luck or through a loophole; broader detection heuristics are an open question.
- Quantify the human verification step. The paper states that experts reviewed a large random sample of corrections but does not report the sample size, the agreement rate, or how many proposals were rejected, leaving the reliability of the semi-automated corrections unmeasured.
- Test the data-quality-over-quantity hypothesis more widely. The comparison rests on a single 3B model against a limited set of baselines, and the paper does not report whether Magma-R1's advantage persists on platforms beyond Android or on the tasks in the curated set that were rewritten rather than merely re-evaluated.
Target Audience
This paper is most useful to researchers and engineers building or evaluating GUI and mobile agents, particularly those training compact vision-language models for on-device deployment. Benchmark designers and data-curation teams will find the purification pipeline and the deficiency taxonomy directly reusable, and the three named deficiency categories give a concrete checklist for auditing their own datasets. Product and strategy readers assessing whether on-device assistants are near-term viable will find the headline cost, size, and success-rate comparisons most relevant. Readers looking for new model architectures will find little here — the contribution is evaluation quality and data curation rather than architectural novelty.
Authors’ abstract
On-device virtual assistants like Siri and Google Assistant are increasingly pivotal, yet their capabilities are hamstrung by a reliance on rigid, developer-dependent APIs. GUI agents offer a powerful, API-independent alternative, but their adoption is hindered by the perception of poor performance, as even the best models (e.g. Qwen3-VL-235B) scores are capped at around 60% on benchmarks like AndroidControl, far from viability for real-world use. Our research reveals that issue lies not only with the models but with the benchmarks themselves. We identified notable shortcomings in AndroidControl, including ambiguities and factual errors, which systematically underrates agent capabilities. To address this critical oversight, we enhanced AndroidControl into AndroidControl-Curated, a refined version of the benchmark improved through a rigorous purification pipeline. On this enhanced benchmark, state-of-the-art models achieve success rates nearing 75% on complex tasks (15% improvement), reflecting that on-device GUI agents are actually closer to practical deployment than previously thought. We introduce our new SOTA model, Magma-R1- 3B, post-trained on just 2.4k curated samples using 60 hours of an H20 GPU (approximately $60). Despite being 200 times smaller in parameters, this model delivers performance comparable to Qwen3- VL-235B. We release both AndroidControl-Curated benchmark and Magma-R1 model to the research community, encouraging adoption of this enhanced benchmark to better reflect model capabilities and accelerate the development of robust, on-device virtual assistants.