Skip to content
AI.info

Research

CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

Overview Research area: Computer vision and computer-use agents, specifically dataset construction and training for interactive CAPTCHA solving in a live browser environment. Technical level: Intermed

CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
arXiv
2609.31957
Published
2026-09-25
Authors
Zhenhao Zhang, Zhaoyu Fan, Haohan Ying, Jingwen Hu, Hancen Fan, Junhao Zhou, Zitian Chen, Linchao Zhu

AI summary

Overview

Research area: Computer vision and computer-use agents, specifically dataset construction and training for interactive CAPTCHA solving in a live browser environment.

Technical level: Intermediate. The paper assumes familiarity with supervised fine-tuning, reinforcement learning (GRPO), LoRA adapters, and vision-language models, but its core arguments are about dataset design and evaluation, which are explained in accessible terms.

Scope: The paper introduces CaptchaArena, a 50K-puzzle, 20-type, 5-interaction-mode training dataset with execution-verified solutions and 46K reasoning-annotated trajectories, and uses it to train CaptchaAgent, a single 9B policy.

What This Paper Is About

Computer-use agents can operate browsers through screenshots and low-level mouse and keyboard actions, but a single failed CAPTCHA can halt an entire workflow. Existing CAPTCHA resources force a trade-off: some cover few challenge types, some provide only answers and click coordinates rather than step-by-step interaction, and some grade clicks against loose geometric regions instead of the true target shape. The paper's goal is to build a large-scale dataset that simultaneously covers many CAPTCHA types, supervises the full screenshot-action loop, and grades irregular targets precisely, then to show that this supervision trains a single policy competitive with much larger closed-source models.

Key Contributions

  1. CaptchaArena. A large-scale, fine-grained computer-use training dataset of 50K interactive CAPTCHA puzzles spanning 20 types and 5 interaction modes (single-click, multi-click, arrow-cycle, real-time, text-entry). Each puzzle has an executable reference solution that is replayed in a real browser and must pass the page's own verifier; puzzle quality is manually inspected. The dataset includes 46K reasoning-annotated trajectories and pixel-mask supervision for irregular targets.

  2. CaptchaAgent. A single 9B policy built on Qwen3.5-9B covering all 20 types, trained with supervised fine-tuning on a subset of the reasoning-annotated trajectories and then with reinforcement learning in which the same environment verifier that validated the dataset solutions directly provides the reward. No learned reward model, preference data, or step-level reward annotation is used.

  3. Per-type error analysis. A breakdown of remaining failures using three diagnostic signals (submit rate, Pass@1, Pass@5), comparing CaptchaAgent with the untrained backbone and with human performance on the same puzzles, organized into three failure modes: non-submission, weak grounding, and unstable execution.

  4. A comparison table positioning existing CAPTCHA resources. Table 1 contrasts prior benchmarks and training datasets on type coverage, direct low-level computer use, multi-step trajectories with intermediate observations, reasoning annotations, RL reward availability, click-evaluation geometry, and scale.

Main Findings

  • Supervised fine-tuning drives most of the gain. SFT raises Pass@1 from 11.4 for the untrained Qwen3.5-9B backbone to 70.5. Reinforcement learning further improves it to 71.7 (+1.2, or +1.5 excluding Click Order), with 16 of 20 types improving.

  • Scale alone does not explain performance. A 35B-A3B Qwen3.5 model also scores 11.4, the same as the 9B backbone. The six open-weight GUI agents evaluated range from 7B to 9B and top out at 35.2, while the three closed-source models (GPT 5.4, Gemini 3.5 Flash, Claude Sonnet 4) span 49.4 to 69.2. CaptchaAgent reaches 70.5 Pass@1 and 86.0 Pass@5 with a 96.8% submit rate; humans reach 94.1 on the same 4,000 puzzles.

  • The untrained backbone's failure is procedural, not perceptual. The 9B backbone averages 56.6 on the 4 types that submit automatically but at most 0.5 on the 16 types requiring an explicit submit action, with an overall submit rate of 20.5%. Training largely removes this: the policy submits on 96.8% of rollouts and on at least 99% in 13 of 20 types.

  • RL gains are largest on mid-difficulty types, not the weakest. Connect Icon improves by +5.1 from 71.7, Place Dot by +4.0 from 16.1, Image Recognition by +3.5 from 56.8, and Coordinates by +3.3 from 96.6. The two weakest types barely move: Dice Count by +1.6 and Patch Select by +0.8. The improvement is larger at Pass@1 (+1.2) than at Pass@5 (+0.8), consistent with RL improving first-rollout reliability more than expanding the solvable set. Click Order is the only substantial regression, dropping 3.1 points.

  • RL also helps on two external benchmarks. Pass@1 rises from 47.2 to 51.0 on Open CaptchaWorld and from 13.6 to 20.0 on Halligan.

  • Four types remain hard after SFT, all with all-or-nothing grading. Place Dot, Patch Select, Slide Puzzle, and Dice Count stay below 45 Pass@1. Patch Select is difficult for every evaluated system: no model exceeds 12.5 while humans reach 73.5. Place Dot is a specific weakness of CaptchaAgent, which submits on 95.3% of rollouts but reaches only 16.1 Pass@1, while GPT-5.4 reaches 88.5 on the same puzzles. Dice Count has a 54.4% submit rate, so nearly half its episodes are scored as failures with the answer never checked; humans reach 83.0 there, their third-lowest score, and the best system evaluated reaches 22.5.

  • Unstable execution dominates the gaps that remain. Pass@1 to Pass@5 rises from 70.5 to 86.0, a 15.5-point gap. The gap exceeds 30 points on Slide Puzzle (42.2 to 91.0), Bingo (56.8 to 92.5), and Pick Area (59.8 to 93.0), but stays under 10 points in the 8 types already above 90.

  • The human gap is concentrated. 8 types reach 90 or higher at Pass@1, 7 of those sit within 5 points of the human reference, and 4 types match or exceed it. Four types account for three-fifths of the 23.6-point gap to the human average of 94.1.

  • Dataset scale as reported. The dataset totals 50K puzzles: 42K training, 4K validation, and 4K test (2,100 training, 200 validation, and 200 test puzzles per type). Table 1 lists CaptchaArena's training scale as 42K. The 3 irregular-target types contribute 7,500 mask-annotated puzzles. SFT uses 600 of each type's 2,100 training puzzles (12,000 total), which per-turn expansion turns into 37,621 training samples; the remaining 1,500 per type are held out for RL.

Methodology in Plain English

The researchers build a self-hosted browser environment rather than scraping live CAPTCHAs. Agents see a fixed 1280×1080 viewport and use five tool calls: screenshot, click(x,y), type_text(s), drag(x0,y0,x1,y1), and timed hold(x,y,Δt). Episodes end when the puzzle is submitted, either by the agent or automatically; submission is automatic in 4 types, and every episode is capped at 15 turns.

Puzzles are generated under task-specific constraints. Geometry Click and Path Finder are procedurally generated in Blender for exact ground truth over geometry, viewpoint, and lighting. Most other images come from ChatGPT Images 2.0, with prompts that diversify object identity, layout, style, and distractors while preserving task semantics. Three types use openly licensed icons and a public reCAPTCHA-V2 image set. Because generators do not always honor the prompt (for example, dice asked to sum to 23 may not), every image is independently inspected by two authors, and the two arithmetic types are manually checked and labeled. Distractors are drawn from the same source as the correct answer, and answer positions are quota-controlled so that fixed guessing strategies gain no systematic advantage.

Answers come from three sources: generation-time labels, manual annotation (double-annotated for Place Dot and Patch Select, with the first author adjudicating), and renderer-derived pixel masks. Each stored answer is compiled into an executable sequence of browser actions, run in the environment with a screenshot recorded before every action, and retained only if the page's verifier accepts the final state. This simultaneously verifies the puzzle and yields one screenshot-action trajectory per puzzle.

A teacher model (GPT-5.4-mini) then receives each turn's screenshot plus the correct action, described in task-level terms rather than as a coordinate, and produces the reasoning that leads to it. An independent judge from a different model family (Gemini-2.5-Flash) checks each generation twice, rejecting on consistency with the verified solution, hindsight phrasing that reveals the answer, and hedging language. Rejected samples return to the teacher with feedback; samples still rejected after three attempts go to a stronger vision model (GPT-5.5), which also checks every Click Order and Connect Icon sample, since weaker models may misread glyphs or miss faint dashed links. Annotation covers the 42K training and 4K validation puzzles, giving 46K annotated trajectories. Test puzzles are never annotated. Each T-turn trajectory is expanded into T per-turn samples because the chat template removes reasoning from previous turns, so every turn is supervised in the same format used at inference.

Training proceeds in two stages. SFT uses rank-64 LoRA adapters applied only to the language model of Qwen3.5-9B, with the vision encoder and vision-language aligner frozen, for 3 epochs; the 200 validation puzzles per type monitor training but do not select checkpoints. RL then continues with GRPO against the live environment, using the verifier as reward. Two design choices matter: difficulty is mined per puzzle rather than per type, because GRPO gives no learning signal when all rollouts for a puzzle receive the same reward, and the reward is monotone both in submitting and in closeness to the answer, so the policy is never rewarded for withholding an uncertain answer.

Evaluation is uniform: every system receives screenshots and predicts pixel-level actions with no agent framework, set-of-mark overlay, or accessibility tree. The test set is all 4,000 puzzles (200 per type); CaptchaAgent is sampled 5 times per puzzle and other systems once, and a puzzle counts as solved only when the server-side verifier accepts the submitted state.

Why This Matters

Impact on research. The paper argues that training data and interaction supervision, not model scale, drive CAPTCHA performance, and it backs that claim with a direct comparison: a 35B-A3B backbone scores the same 11.4 as a 9B backbone, while a fine-tuned 9B policy reaches 70.5. It also contributes a reusable supervision recipe whose reward signal requires no learned reward model and no human reward labels, and it upgrades click grading from bounding boxes or manually defined regions to pixel masks for irregular targets, addressing a case where region-based grading can over-accept areas outside the true target.

Real-world applications:

  • Long-horizon web automation and robotic process automation, where a single unsolved CAPTCHA currently blocks everything downstream in a workflow.
  • Benchmarking and stress-testing general GUI agents, since a CAPTCHA requires the full loop of reading the page, deciding, acting on the right pixels, and verifying the state change.
  • Accessibility tooling for users who cannot complete interactive challenges themselves.
  • Defensive CAPTCHA research, since the paper releases the dataset to support study of CAPTCHA robustness as well as agent evaluation.

Industry relevance. Vendors building browser agents and computer-use products now have a public, execution-verified training set and a single 9B reference policy that outperforms much larger closed-source models on this task (71.7 versus 49.4 to 69.2 for GPT 5.4, Gemini 3.5 Flash, and Claude Sonnet 4). The capability is explicitly framed as dual use, with the environment fully self-hosted and no proprietary CAPTCHA assets scraped or redistributed.

Future Directions

  • Closing the precise-grounding gap. Place Dot is the clearest residual weakness: CaptchaAgent submits on 95.3% of rollouts but reaches only 16.1 Pass@1, against 88.5 for GPT-5.4 and 100.0 for humans. The authors also flag the frozen vision encoder and this task as their largest gap to the human reference.
  • Reducing execution variance. The 15.5-point Pass@1-to-Pass@5 gap, exceeding 30 points on Slide Puzzle, Bingo, and Pick Area, suggests a valid solution is often reachable but unreliable on the first rollout.
  • Improving all-or-nothing task types. The four hardest types (Place Dot, Patch Select, Slide Puzzle, Dice Count) all grade with a single-predicate check, and Dice Count's 54.4% submit rate means many failures occur before the answer is ever checked.
  • Reducing dependence on commercial models for annotation. The limitations section notes that reasoning annotations depend on commercial models, though they explain only replay-verified actions.
  • Modeling mouse trajectories. Because the action space is discrete, the released trajectories cannot reproduce mouse-trajectory signals used by behavioral detectors, closing off a dimension of realism.

Target Audience

Researchers and engineers working on computer-use and GUI agents, dataset builders interested in execution-verified supervision and reward design, and practitioners building web automation pipelines that must survive interactive challenges. It is also relevant to people studying CAPTCHA robustness and human-versus-agent capability gaps, since it reports a direct human reference of 94.1 on the same 4,000 puzzles used to evaluate every system.

Authors’ abstract

Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving. It contains 50K puzzles across 20 CAPTCHA types and 5 interaction modes, with every solution verified through execution. CaptchaArena provides 50K screenshot-action trajectories, including 46K with step-by-step reasoning annotations. It also includes fine-grained pixel-mask annotations for irregular targets. Using CaptchaArena, we train CaptchaAgent, a single 9B policy for all 20 CAPTCHA types, with supervised fine-tuning followed by reinforcement learning. The environment verifier directly provides the RL reward. Supervised fine-tuning reaches 70.5 Pass@1, and reinforcement learning further improves it to 71.7, while also improving performance on two external benchmarks. These results demonstrate the value of large-scale, fine-grained computer-use supervision for training interactive CAPTCHA agents. We release CaptchaArena and CaptchaAgent at https://github.com/X0X0X00/CaptchaArena.

Read the original paper