Skip to content
AI.info

Research

DigiData: Training and Evaluating General-Purpose Mobile Control Agents

Overview Research area: AI agents that operate mobile devices ("computer use" / mobile control agents), with a focus on dataset construction, benchmark design, and LLM-based evaluation. Technical leve

DigiData: Training and Evaluating General-Purpose Mobile Control Agents
arXiv
2511.07413
Published
2025-11-10
Authors
Yuxuan Sun, Manchen Wang, Shengyi Qian, William R. Wong, Eric Gan, Pierluca D'Oro, Alejandro Castillejo Munoz, Sneha Silwal, Pedro Matias, Nitin Kamra, Satwik Kottur, Nick Raines, Xuanyi Zhao, Joy Chen, Joseph Greer, Andrea Madotto, Allen Bolourchi, James Valori, Kevin Carlberg, Karl Ridgeway, Joseph Tighe

AI summary

Overview

Research area: AI agents that operate mobile devices ("computer use" / mobile control agents), with a focus on dataset construction, benchmark design, and LLM-based evaluation.

Technical level: Intermediate. The paper is readable without deep technical background, but familiarity with supervised fine-tuning, multimodal models, and agent evaluation metrics helps.

Scope in one sentence: The paper introduces DigiData, a 152,000-trajectory mobile control dataset covering 8,275 goals across 26 Android apps, together with DigiData-Bench, a 309-goal benchmark across 37 apps that supports human-assisted and AI-assisted dynamic evaluation, and it uses both to train and measure mobile control agents.

What This Paper Is About

Training an AI agent to operate a phone—tapping, scrolling, and navigating apps to complete a user's goal—requires two things that have been missing: data that covers how apps actually work, and evaluation methods that reliably say whether a goal was achieved. Existing mobile control datasets mostly derive goals from unstructured interactions, so they under-cover advanced app features, and existing evaluation largely relies on "step accuracy," which only checks whether each predicted action matches one recorded human action. The paper's goal is to build a dataset whose goals come from deliberate, exhaustive exploration of app features, and to provide a benchmark and evaluation protocols that measure whether an agent actually completes a real-world task.

Key Contributions

  1. DigiData, a large-scale mobile control dataset. 152,000 trajectories and 8,275 unique goals across 26 Android apps, collected through a three-phase pipeline (goal curation, demonstrations collection, trajectory verification), with an average of 9.2 steps per goal and 318 goals per app.

  2. A multi-modal dataset. Each step includes a screenshot, a UI Tree from the Android accessibility tree, and Chain-of-Thought annotations generated with Llama 4 (screen description, action description, rationale, expected UI change).

  3. DigiData-Bench, a benchmark with rigorous evaluation protocols. 309 goals across 37 Android apps and 8 app categories, split into Seen, Familiar, and Novel app groups, supporting human-assisted dynamic evaluation and an automated subset (DigiData-Bench-Auto) using LLM judges on a tethered Android device.

  4. The first empirical evaluation of LLM judges for measuring mobile control agent success. Including a fine-tuned Llama 4 judge trained on 7,057 failed and 2,034 successful model-generated trajectories, compared against GPT4o and step accuracy as an evaluation signal.

Main Findings

  • Goal depth and diversity exceed prior datasets. DigiData goals average 9.2 steps, roughly 50% or more than existing datasets, and it averages 318 goals per app versus 85 for AitW and 17 for AndroidControl. Its goal diversity score is 0.45, compared to 0.43 for AndroidControl and 0.35 for AitW.

  • Data quality is higher than AitW by human verification. When 20k randomly sampled AitW trajectories were passed through the authors' human verification system, only 84% achieved their prescribed goal, versus 94.6% of DigiData before verification and 100% after verification. The verification procedure filters out 5.3% of raw trajectories.

  • Step accuracy is a poor proxy for agent quality. In Figure 5 (left), rankings by step accuracy do not reliably match human-evaluated task success. The 3B and 8B models have nearly identical step accuracy on DigiData-Bench and AitW but significantly different success rates on DigiData-Bench. Step accuracy's Kendall rank correlation with human judgment is 0.72, well below the LLM judges.

  • Chain-of-Thought training improves both performance and explainability. The 8B COT model reaches 47.3% success rate on DigiData-Bench (human-evaluated), versus 42.1% for the 8B model without COT, and 53.6% versus 48.5% under the GPT4o LLM judge.

  • Agents trained with DigiData beat strong zero-shot baselines. The 8B COT model gets 47.3% on all DigiData-Bench goals versus 27.8% for GPT4o and 39.2% for Qwen2.5VL. On the Novel app subset the 8B COT model reaches 36.7%, compared to 26.5% for GPT4o and 40.8% for Qwen2.5VL.

  • Human evaluation is low-variance, and humans remain far ahead. Running human evaluation on the 3B model three times gave a standard deviation of ±0.4%, and human expert performance was 90.1% task success rate.

  • LLM judges correlate well with human judgment. GPT4o as a judge reached 0.89 accuracy, 0.87 precision, 0.86 recall, 0.9 NPV, 0.9 TNR, and 0.94 rank correlation. A fine-tuned Llama 4 Scout reached 0.87 accuracy, 0.88 precision, 0.79 recall, 0.86 NPV, 0.92 TNR, and 0.89 rank correlation—slightly outperforming GPT4o on TNR, the key metric for flagging unfit trajectories during data collection.

  • Data scaling helps seen and familiar apps but not novel apps. Training four models with varying data volumes showed overall performance improving with more data, but no significant improvement on novel apps, pointing to limits of supervised fine-tuning and motivating reinforcement learning.

  • Performance drops with task complexity. A preliminary analysis using average step count as a proxy for complexity found that success rates tend to decrease as the number of steps increases.

Methodology in Plain English

The dataset was built by humans, not scraped from existing logs. For each of 26 apps, trained annotators were instructed to exhaustively explore the app's screens and menus and write down goals that cover its features—including deeper features, which were sometimes broken into sequences of goals to keep any single goal from becoming too complex. Those goals were then demonstrated on a custom Android app deployed on physical or emulated devices; the app selected a goal, recorded the screen state and the annotator's actions, and let the annotator review the trajectory before upload. Finally, every trajectory was checked: an LLM judge tuned to have a low false accept rate scored it, and trajectories marked unsuccessful were sent to human annotators, being kept if either judge called them successful.

The Chain-of-Thought annotations were generated afterwards by Llama 4, producing the observation, action, thought, and expected UI change for each step.

For evaluation, the authors built DigiData-Bench and designed two protocols. In human-assisted dynamic evaluation, a human worker follows goal-specific state-initialization instructions (so the goal is meaningful and reproducible), monitors the agent, blocks unsafe actions, and judges success when the trajectory ends. In AI-assisted evaluation (DigiData-Bench-Auto), roughly half the goals—those needing no manual setup, login, or location and with low risk—are run end-to-end with apps reset automatically and an LLM judge scoring success from the screenshot and UI tree. A per-step action-matching test set provides the offline "step accuracy" metric for comparison.

The agent itself is a fine-tuned Perception Language Model (PLM) trained with supervised fine-tuning on a mixture of DigiData, AitW, AndroidControl (low-level and high-level), and Cauldron (a general VQA dataset, included to prevent catastrophic forgetting). Training used 128 NVIDIA A100 GPUs with 32 more for evaluation, a learning rate of 2e-5, a batch size of 2 samples per GPU, 60k iterations, and image tiles of 448x448 with at most 4 tiles per image.

Why This Matters

Impact on research. The paper argues that step accuracy—the dominant offline metric in mobile control—can rank models incorrectly, and supplies evidence that dynamic, trajectory-level success evaluation plus LLM judges gives a stronger signal. It also releases multi-modal data (screenshots, UI trees, Chain-of-Thought) at a scale not previously available with permissive licensing, which opens observation-encoding and explainability research that smaller datasets could not support.

Real-world applications:

  • Accessibility tools that operate a phone on behalf of users who cannot navigate touch interfaces.
  • Automating multi-step, app-specific chores such as organizing photos into new folders or configuring shipping locations.
  • Helping users reach advanced app features that are unfamiliar and often unavailable through public APIs, since the paper notes these are frequently interface-only.
  • Automated regression testing of mobile apps by agents that carry out real user goals.

Industry relevance. The authors are from FAIR at Meta and Meta Reality Labs, and the work targets general-purpose device control—an area where both proprietary models (GPT4o) and open models (Qwen2.5VL) are already being applied. The result that a fine-tuned open-weight Llama 4 judge approaches GPT4o's judging accuracy on a few thousand samples is directly relevant to teams who want automated, low-cost evaluation without sending data to a proprietary service.

Future Directions

  • Scope beyond 26 apps. The paper states DigiData is limited to 26 mobile apps, which may not match the diversity of real-world applications, and calls for expanding to a broader range of applications.
  • Generalization to novel apps. Performance did not significantly improve on novel apps as data scaled, which the authors attribute to limitations of supervised fine-tuning and suggest addressing with reinforcement learning.
  • Transfer to non-mobile interfaces. The authors note that generalizability to more general computer-based interfaces remains unexplored.
  • Reducing evaluation bias. The paper acknowledges that its human-assisted and AI-assisted protocols may introduce biases or limitations, motivating further work on automated evaluation quality.

Target Audience

Researchers and engineers building GUI or device control agents, teams constructing agent training datasets and benchmarks, and practitioners who need reliable evaluation of agents in the wild rather than offline action-matching scores. It is also relevant to those studying LLM-as-judge methodology, since the paper provides a rare calibrated comparison of judge models against human success judgment. The paper contains a large amount of app and goal metadata, so product teams evaluating whether an agent can handle their own app category will also find it useful.

Authors’ abstract

AI agents capable of controlling user interfaces have the potential to transform human interaction with digital devices. To accelerate this transformation, two fundamental building blocks are essential: high-quality datasets that enable agents to achieve complex and human-relevant goals, and robust evaluation methods that allow researchers and practitioners to rapidly enhance agent performance. In this paper, we introduce DigiData, a large-scale, high-quality, diverse, multi-modal dataset designed for training mobile control agents. Unlike existing datasets, which derive goals from unstructured interactions, DigiData is meticulously constructed through comprehensive exploration of app features, resulting in greater diversity and higher goal complexity. Additionally, we present DigiData-Bench, a benchmark for evaluating mobile control agents on real-world complex tasks. We demonstrate that the commonly used step-accuracy metric falls short in reliably assessing mobile control agents and, to address this, we propose dynamic evaluation protocols and AI-powered evaluations as rigorous alternatives for agent assessment. Our contributions aim to significantly advance the development of mobile control agents, paving the way for more intuitive and effective human-device interactions.

Read the original paper