Research
"Are We Done Yet?": A Vision-Based Judge for Autonomous Task Completion of Computer Use Agents
Overview Research area: Artificial intelligence, specifically Computer Use Agents (CUAs) and Vision–Language Model (VLM) based evaluation of GUI task completion. Technical level: Intermediate. The met

- arXiv
- 2511.20067
- Published
- 2025-11-25
- Authors
- Marta Sumyk, Oleksandr Kosovan
AI summary
Overview
- Research area: Artificial intelligence, specifically Computer Use Agents (CUAs) and Vision–Language Model (VLM) based evaluation of GUI task completion.
- Technical level: Intermediate. The methods are conceptually accessible (screenshot plus task description in, judgment out), but the paper assumes familiarity with CUAs, VLMs, and benchmark evaluation.
- Scope in one sentence: The paper builds a vision-based "judge" that reads the final screenshot of a macOS desktop and decides whether a CUA finished its task, then feeds that judgment back to the agent so it can retry.
This work was accepted to appear at the AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent).
What This Paper Is About
Computer Use Agents can click, type, and navigate real interfaces, but they cannot reliably tell whether they actually finished the job — they either claim success when the task is unfinished, or finish successfully and keep acting redundantly. The paper's goal is to replace hand-written verification scripts with a zero-shot VLM judge that looks only at the final screenshot and the task description, produces a done/not done verdict with a short written rationale, and passes that rationale back to the agent for another attempt. The authors build this for macOS, an environment they describe as underexplored relative to web and mobile.
Key Contributions
- A human-labeled dataset of 1,260 tasks across 42 built-in macOS applications (30 tasks per application), published for reproducible evaluation.
- A zero-shot VLM-based methodology for autonomously evaluating task completion, reporting up to 73% classification accuracy against human-annotated ground truth.
- An evaluation-feedback pipeline in which the VLM's natural-language rationale is returned to the CUA, which retries from its current state rather than restarting, yielding an average relative improvement of 27% in overall task success rate.
- A cross-model comparison of five VLM evaluators — GPT-4o, Claude 3.5 Sonnet, LLaVA-v1.5-7B, InternVL 2-8B, and Qwen2-VL-7B — spanning proprietary and open-source families, evaluated against three CUAs: Claude Computer Use, OpenAI Operator, and UI-TARS.
Main Findings
- Highest classification accuracy: 73%. Claude 3.5 Sonnet reached 0.73 accuracy on UI-TARS, the top proprietary evaluator result. It also scored 0.69 on OpenAI Operator and 0.71 on Anthropic CU.
- GPT-4o results: 0.61 on OpenAI Operator, 0.69 on Anthropic CU, and 0.64 on UI-TARS.
- Best open-source evaluator: Qwen2-VL-7B, with 0.68 on OpenAI Operator, 0.66 on Anthropic CU, and 0.70 on UI-TARS.
- Weakest open-source evaluator: LLaVA-v1.5-7B, scoring 0.56 on OpenAI Operator, 0.61 on Anthropic CU, and 0.52 on UI-TARS.
- InternVL 2-8B results: 0.62 on OpenAI Operator, 0.67 on Anthropic CU, and 0.61 on UI-TARS.
- Feedback improves success rates. Every evaluated VLM feedback mechanism produced measurable gains over the no-feedback baseline after only one retry.
- Proprietary evaluators give the largest gains, achieving up to 61% relative success rate improvements; open-source evaluators such as Qwen2-VL-7B provided consistent boosts as well.
- Weaker agents benefit most. Agents with a lower baseline success rate, such as Anthropic CU, improved the most from visual feedback.
- Overall headline numbers: up to 73% classification accuracy in task success detection and an average relative improvement of 27% in overall task success rate (the contributions section phrases this as 27% relative percentage points on average; the abstract and conclusion phrase it as a 27% relative improvement).
- Dataset scale in context: 1,260 tasks, compared with 369 tasks in the OSWorld benchmark.
Methodology in Plain English
The pipeline has three steps. First, a CUA is given a task description and attempts it inside macOS; the full trajectory is recorded, including step-by-step screenshots, the actions taken (clicks, double-clicks, typed text, key presses, and waits), and the agent's reasoning at each step.
Second, a VLM receives only the final screenshot and the original task description — no logs, no internal agent state — and is prompted in a zero-shot setting to output a binary judgment (done/not done) plus a short natural-language rationale for the decision. Five VLMs are tested in this role. Because the evaluator is a separate model from the acting agent and sees only the screen, the judgment is meant to be free of the acting agent's internal assumptions.
Third, if the verdict is not done, the rationale is handed back to the CUA as feedback. The agent replans and retries, resuming from its current state rather than restarting the whole trajectory. This is intended to reduce both outright task failure and redundant actions.
The dataset was designed to cover productivity, communication, multimedia, system utilities, and developer tools, with tasks ranging from simple ("Open Calendar app") to multi-step ("Filter apps by free in App Store and open the first result"). The authors deliberately avoided tasks requiring private user data or external configuration, such as importing files or logging into accounts. The three CUAs were chosen because they currently achieve leading performance on OSWorld.
Why This Matters
Impact on research. The paper argues that CUAs rely purely on visual observation and can therefore fail silently or partially when interfaces are unexpected, occluded, or shifted. Existing evaluation either depends on manually written verification scripts (as in OSWorld) that break when interfaces change, or lives in browser and simulated environments where success states are explicitly defined. This work extends autonomous evaluation to unstructured desktop interfaces with no universal representation like an HTML tree, complementing earlier work such as AutoEval in robotics (which reported a 99% reduction in human annotation time).
Real-world applications:
- Desktop automation assistants that need to know when to stop acting, avoiding wasted compute on already-finished tasks.
- Verification layers for enterprise workflows on macOS, where auditability of agent actions matters.
- Self-correction loops that let agents recover from partial or failed task attempts without human intervention.
- A reusable judge that could serve as a reward signal for training agents, reducing dependence on large human-labeled datasets.
Industry relevance. Companies deploying CUAs must trust that an agent's "task complete" claim is true; false completion claims undermine user trust, while unrecognized completion wastes computation. A model-agnostic, screenshot-based judge needs no OS-specific instrumentation and could in principle transfer across platforms, which matters for vendors shipping agents across operating systems. The paper also notes that reliable evaluation is fundamental for measuring performance and enabling self-improvement of agents generally.
Future Directions
- Cross-platform expansion. Extend beyond macOS to Linux and Windows, since windows, menus, buttons, and text are common across all three, and the vision-based approach needs no OS-specific instrumentation.
- Step-level evaluation. Move from the current binary, end-of-task metric to judging each intermediate action by whether it moves the agent closer to the goal, plus an ablation study on how many screenshots or temporal observations are most informative.
- Evaluator robustness analysis. Inter-model agreement analysis, calibration measurements, and consistency studies across VLMs, using techniques such as temperature scaling, conformal prediction, or ensemble averaging to provide confidence intervals for success predictions.
- Reward signal and multi-agent integration. Use the evaluator's output directly as a reward in reinforcement learning pipelines for CUAs, and extend toward multi-agent frameworks where an evaluator continuously monitors actions and delivers real-time feedback on each step.
Target Audience
Researchers and engineers working on GUI agents, computer-use agents, and multimodal evaluation; practitioners who deploy desktop automation and need reliable completion signals; and benchmark designers interested in replacing script-based verification. Readers tracking agent reliability and trustworthiness will find the evaluator-accuracy and feedback-loop results most directly useful, while those interested in dataset construction for agent evaluation will benefit from the 1,260-task, 42-application macOS dataset, available at the Zenodo record cited in the paper, with code at the linked GitHub repository.
Authors’ abstract
Computer Use Agents (CUAs) are designed to autonomously operate digital interfaces, yet they often fail to reliably determine whether a given task has been completed. We present an autonomous evaluation and feedback framework that uses vision-language models to assess task completion directly from screenshots and task descriptions. Our dataset covers 42 built-in macOS applications and 1,260 human-labeled tasks across a wide range of scenarios. Our framework achieves up to 73 percent accuracy in task success detection and yields an average relative improvement of 27 percent in overall task success when evaluator feedback is applied. These results show that vision-based evaluation can serve as an effective feedback mechanism that improves the reliability and self-correction of autonomous computer-use agents.