Skip to content
AI.info

Research

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Overview Research area: Computer-use agents (CUAs) that operate graphical user interfaces, online reinforcement learning, and on-policy self-distillation. Technical level: Advanced (assumes familiarit

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
arXiv
2609.40253
Published
2026-09-30
Authors
Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen

AI summary

Overview

  • Research area: Computer-use agents (CUAs) that operate graphical user interfaces, online reinforcement learning, and on-policy self-distillation.
  • Technical level: Advanced (assumes familiarity with policy optimization, token-level log-probabilities, and agent training loops).
  • Scope: The paper proposes ComputerSD, a method that turns per-step feedback from executed GUI actions into value-gated token-level supervision, combined with trajectory-level GRPO, for training computer-use agents online.

What This Paper Is About

Computer-use agents trained online usually receive only a sparse reward at the end of a task, which says whether the task succeeded but not which of the many intermediate actions helped or hurt. On-policy self-distillation (OPSD) can supply token-level credit, but the authors argue it fails in CUA training for two reasons: privileged guidance fixed before the rollout becomes misaligned with the state the student actually reaches, and the token-level probability shifts it induces often contradict whether the step was correct. The goal of this work is to derive both the guidance and a step-level judgment from the real-time observation after each action, and to use that judgment to regulate the token-level signals.

Key Contributions

  1. ComputerSD method. An online self-distillation method in which a fine-tuned GUI analyzer writes real-time feedback after each executed action, and a value gate weights each token-level OPSD signal by its agreement with the analyzer's judgment of that step.
  2. Value-gated OPSD formulation. The guidance from the analyzer is added as privileged context for rescoring, and the step-level value score scales the resulting log-probability gap through a sigmoid gate (Equation 6 and 7), combined with trajectory-level GRPO under a coefficient λ_OPSD (Equation 8).
  3. Fully asynchronous training pipeline. Environment interaction, GUI analysis, privileged rescoring, and policy optimization are overlapped so rollout workers keep collecting trajectories while each executed step is analyzed and rescored.
  4. Empirical demonstration on OSWorld-Verified. With two 8B backbones of differing computer-use specialization (Qwen3-VL-8B-Thinking and EvoCUA-8B), ComputerSD outperforms outcome-only GRPO, and removing either real-time feedback or the value gate drops performance below GRPO.

Main Findings

  • OSWorld-Verified improvement: ComputerSD raises success rate from 37.9% to 39.8% on Qwen3-VL-8B-Thinking (a 1.9 point gain over GRPO) and from 43.8% to 47.9% on EvoCUA-8B (a 4.1 point gain).
  • Out-of-distribution generalization: On categories held out from training, ComputerSD reaches 27.5% versus 21.6% for GRPO on Qwen3-VL-8B-Thinking, and 35.4% versus 30.2% on EvoCUA-8B. On WindowsAgentArena, ComputerSD scores 24.5 versus 23.3 for GRPO on Qwen3-VL-8B-Thinking, and 27.8 versus 24.2 on EvoCUA-8B.
  • In-domain gains: In-domain OSWorld-Verified results are 47.5 for ComputerSD versus 48.2 for GRPO on Qwen3-VL-8B-Thinking, and 55.7 versus 52.2 on EvoCUA-8B.
  • GRPO degrades OOD while ComputerSD does not: The paper reports that outcome-only GRPO improves in-domain performance but degrades most OOD performance relative to base models on both backbones — 21.6 (−1.4) versus a 23.0 base, and 30.2 (−1.4) versus a 31.6 base.
  • Ablations isolate two failures: Removing GUI analyzer SFT lowers success to 38.0% (−1.8 points); replacing real-time feedback with fixed guidance lowers it to 35.0% (−4.8 points); removing the value gate lowers it to 36.6% (−3.2 points), all below the 39.8% full configuration.
  • Gate design matters: The SDAR gate yields 38.1%, a hard binary value gate 35.2%, and a reverse value gate 36.0%, compared with 39.8% for the proposed gate.
  • OPSD coefficient sensitivity: λ_OPSD of 0.1 gives 35.6%, 0.01 gives 39.8%, and 0.001 gives 37.1%. At 0.1 the KL loss rises by more than an order of magnitude relative to the other settings.
  • Faster learning per trajectory: ComputerSD reaches the final GRPO performance after approximately 100 policy updates on Qwen3-VL-8B-Thinking and 40 on EvoCUA-8B, using about 56% and 22% of GRPO's trajectory budget respectively.
  • Throughput cost and recovery: Asynchronous GRPO and ComputerSD stabilize at approximately 370 and 210 trajectories per hour. ComputerSD retains about 57% of outcome-only GRPO's throughput, and the asynchronous framework achieves five times the training throughput of synchronous ComputerSD.
  • Motivating measurement: The paper reports that over 80% of tokens at correct steps are suppressed and nearly 20% of tokens at incorrect ones are reinforced throughout training under OPSD guidance.

Methodology in Plain English

The approach starts by collecting a labeled dataset for a helper model. The authors run a base policy on 222 training tasks, sampling eight trajectories per task, which yields 1,776 trajectories and 34,188 step-level samples. A strong expert model, Kimi K3, then annotates each step with a value score and forward-looking guidance; 43.4% of trajectories succeeded and 56.6% failed, while 40.1% of steps were judged correct and 59.9% incorrect. A GUI analyzer initialized from Qwen3-VL-8B-Thinking is fine-tuned with LoRA on these annotations and then frozen.

During online training, after each action the analyzer reads the pre- and post-action screenshots and returns a step-level value score plus guidance. That guidance is appended to the agent's ordinary context as privileged information, and the sampled response is rescored under both contexts. The difference in log-probabilities is the token-level signal. The value score scales it, and a sigmoid gate turns the scaled value into a per-token weight, so signals consistent with the step judgment are reinforced and inconsistent ones attenuated.

This token-level objective is added to standard trajectory-level GRPO with a small coefficient, so step feedback refines credit assignment within an episode while the terminal reward keeps learning anchored to task success. To keep the per-step analysis from stalling rollout, environment interaction, analysis, rescoring, and optimization run asynchronously, and updated parameters are published to rollout workers as soon as each training step finishes. At evaluation time, only the ordinary context is used, with no analyzer calls or guidance.

Why This Matters

  • Impact on research: The work argues that for multi-turn GUI agents, the useful place to intervene in self-distillation is the level of the individual step, not the teacher–student gap or the trajectory outcome. It also gives a concrete recipe for making OPSD usable in online CUA training, where sparse outcome rewards leave credit assignment unresolved.
  • Real-world applications:
    • Training desktop automation agents that complete multi-step tasks across office applications.
    • Improving agents for settings where OOD robustness matters, since ComputerSD improved held-out application categories and the cross-platform WindowsAgentArena benchmark where GRPO regressed.
    • Cost-effective agent improvement, since a fine-tuned 8B analyzer replaces repeated calls to a large expert model during training.
    • Post-training specialized computer-use models further, since gains appeared on the already fine-tuned EvoCUA-8B backbone.
  • Industry relevance: The reported fivefold throughput gain from asynchronous execution, and the finding that ComputerSD retains about 57% of GRPO's throughput, speak directly to the practical cost of deploying online training loops for GUI agents. The listed comparison models include systems such as Claude-4.5-Sonnet (62.9), UI-TARS-2 (47.5), and Claude-4-Sonnet (43.9), placing an 8B model trained with ComputerSD at 47.9 in the same range as much larger proprietary systems for this benchmark.

Future Directions

  • Analyzer reliability: The limitations section notes that ComputerSD depends on the GUI analyzer producing both reliable guidance and reliable step judgments. When both are unreliable, the value gate cannot fully correct the signal, and systematic analyzer errors may still bias policy updates.
  • Reducing feedback cost without losing reliability: The authors state that improving feedback reliability without substantially increasing inference cost remains an important direction.
  • Extending evaluation: The paper evaluates on OSWorld-Verified and WindowsAgentArena; whether the approach transfers to other GUI environments and longer interaction horizons is left open, though the authors note a maximum of 30 interaction steps during training and 50 during evaluation.
  • Analyzer quality measurement: The appendix reports a held-out validation set of 341 samples (1% of annotated samples) for comparing the base model with the fine-tuned analyzer, but the reported numbers are not included in the supplied content.

Target Audience

Researchers and engineers working on GUI agents, computer-use agents, and agentic reinforcement learning, particularly those interested in online training, credit assignment over long action sequences, and self-distillation methods. It is also relevant to practitioners who need to make online agent training affordable, given the asynchronous pipeline and throughput measurements, and to readers who want an ablation-level view of why step-level gating and real-time guidance matter relative to outcome-only GRPO.

Authors’ abstract

Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.

Read the original paper