Research
DrawingBench: Evaluating Spatial Reasoning and UI Interaction Capabilities of Large Language Models through Mouse-Based Drawing Tasks
DrawingBench: Evaluating Spatial Reasoning and UI Interaction Capabilities of Large Language Models through Mouse-Based Drawing Tasks Overview Research area: NLP / evaluation of agentic large language

- arXiv
- 2512.01174
- Published
- 2025-12-01
- Authors
- Hyunjun Kim, Sooyoung Ryu
AI summary
DrawingBench: Evaluating Spatial Reasoning and UI Interaction Capabilities of Large Language Models through Mouse-Based Drawing TasksOverview
Research area: NLP / evaluation of agentic large language models; spatial reasoning; GUI action generation and benchmark design.
Technical level: Intermediate. The concepts (benchmarks, prompts, multi-turn feedback, agentic UI control) are approachable, but the paper assumes some familiarity with LLM evaluation practice and agent architectures.
Scope in one sentence: The paper introduces DrawingBench, a transparent and reproducible benchmark that scores LLMs on how well they can emit low-level mouse action sequences to draw specified shapes on a canvas, and evaluates four frontier models over 1,000 tests.
What This Paper Is About
Existing benchmarks for spatial reasoning (such as CLEVR, RAVEN, and SpartQA) ask models static questions, and existing UI benchmarks (such as WebArena and MiniWoB++) rely on high-level actions like clicking links rather than precise coordinate-level control. Neither probes whether a model can produce a spatial configuration through a sequence of executable GUI actions, and most evaluations are opaque and hard to audit. DrawingBench addresses this by giving models natural-language drawing prompts and requiring them to output low-level mouse action sequences (move, click, mouseDown, mouseUp) that an automated rule-based system then evaluates against eight objective criteria.
Key Contributions
- DrawingBench, described as the first benchmark integrating spatial reasoning with GUI action execution, with 250 prompts across 20 categories and four difficulty levels.
- Automated evaluation built on eight quantitative criteria, four error types, and reproducible scoring.
- A multi-turn protocol with structured external feedback that tests correction from explicit, rule-generated signals rather than self-critique.
- An empirical study of 1,000 trials with four state-of-the-art LLMs (Claude-4 Sonnet, GPT-4.1, GPT-4.1-mini, Gemini-2.5 Flash) analyzing performance across difficulty levels, plus an open-source release of code and data.
Main Findings
- High overall performance: Models averaged 0.925 in Turn 1 and 0.954 in Turn 2 (+0.030). Perfect-score rate rose from 76.8% (192/250) to 92.8% (232/250), an increase of 40 perfect scores. Median score moved from 0.970 to 0.985, and standard deviation fell from 0.099 to 0.066 (the paper describes this as a 33% reduction in variance).
- Cross-model consistency: All four models exceeded 90% perfect rate in Turn 2: Claude-4 Sonnet 94.4%, GPT-4.1 93.2%, Gemini-2.5 Flash 92.0%, GPT-4.1-mini 90.8%. Turn 1 scores were Claude-4 Sonnet 0.931, GPT-4.1 0.927, Gemini-2.5 Flash 0.923, GPT-4.1-mini 0.919.
- The "difficulty paradox": Hard tasks scored highest in Turn 2 (0.973), above Medium (0.958), and achieved a 100.0% perfect rate versus Medium's 92.8%. The paper attributes this to specification clarity: Hard tasks averaged 4.2 explicit constraints versus Medium's 2.8, and 100% of Hard tasks had position constraints versus 62% of Medium tasks.
- Very Hard tasks remain the weak point: Turn 2 score 0.869 with a 60.0% perfect rate, though they also showed the largest absolute gain (+0.052). These 20 tasks produced 35% of all Turn 1 errors (24 of 68), an error rate of 1.2 errors per task versus 0.18 for Easy tasks.
- Structured feedback works, unevenly: 71.2% of tasks needed no second turn (early stopping when Turn 1 score ≥ 0.9). Of the rest, 15.2% improved by 0.01–0.05, 8.8% by 0.06–0.10, and 4.8% by more than 0.10. The scenes category improved the most (+0.328, from 0.672 to 1.000), followed by angle (+0.100) and creative (+0.076).
- Errors fell substantially under feedback: Total errors dropped from 68 to 35 (-49%). SYNTAX_ERROR went 3 to 0 (-100%), COORDINATE_ERROR 8 to 2 (-75%), LOGIC_ERROR 15 to 5 (-67%), and EFFICIENCY_WARNING 42 to 28 (-33%).
- Error patterns cluster around specification ambiguity: Tool selection omissions occurred in 18 tasks, and coverage insufficiency in 12 tasks, mostly when tools or sizes were implied rather than stated. The discussion reports tool selection errors at 15%, coordinate precision errors at 10%, and coverage insufficiency at 8%.
- Action behavior: Average action sequence length was 15.3 actions in Turn 1 and 16.2 in Turn 2. Drawing actions (mouseDown/mouseUp pairs) rose from 42.5% to 45.8%, with drawing segments going from 6.5 to 7.4. Pen was the most-used tool (68.2%), followed by rectangle (18.4%), circle (12.7%), line (8.5%), fill (5.2%), and eraser (2.1%).
Methodology in Plain English
The team built a browser-based drawing application with a 1000 × 700 pixel canvas positioned at screen coordinates (90, 70). It offers six tools (pen, eraser, fill, line, rectangle, circle), eight colors (black, red, green, blue, yellow, magenta, cyan, white, each mapped to a hex code), and three brush sizes (2px, 5px, 10px). Models never see the canvas; they receive only a text prompt and must output a JSON list of four possible actions: moveTo(x, y), click(), mouseDown(), and mouseUp().
Prompts are organized into four difficulty levels and 20 categories, with 250 prompts total. Each model attempt is scored by a deterministic rule engine against eight weighted criteria: required tools (0.20), required colors (0.20), minimum segments (0.15), canvas coverage (0.10), position constraint (0.15), size constraint (0.10), syntax validity (0.05), and coordinate bounds (0.05). Errors carry fixed penalties: SYNTAX_ERROR -0.3, COORDINATE_ERROR -0.2, LOGIC_ERROR -0.1, EFFICIENCY_WARNING -0.05.
The setup runs at most two turns. After Turn 1 the system produces structured feedback consisting of the current score and grade, specific error descriptions, actionable advice, and missing criteria; the model then revises. If the Turn 1 score is at least 0.9, Turn 2 is skipped. The authors report that pilot tests showed diminishing returns after Turn 2 (under 1% improvement), and that early stopping applies to 71.2% of tasks. Testing used temperature 0.7, a 4,000-token cap, a 30-second timeout, and up to 3 retries, with all 1,000 tests (250 prompts × 4 models) run sequentially over 5.6 hours. Per-model wall-clock times were 77.1 minutes for GPT-4.1-mini (fastest), 84.2 for Claude-4 Sonnet, 86.3 for GPT-4.1, and 90.6 for Gemini-2.5 Flash. Average token usage per test was 1,246 input and 388 output tokens (1,634 total).
Why This Matters
The paper argues that as agentic AI systems act autonomously, trust depends on evaluations that stakeholders can inspect and reproduce. DrawingBench's eight criteria produce deterministic 0–1 checks and its action-level records let an auditor trace a score of, say, 0.85 to a specific failure such as "required tool not used," rather than treating the model as a black box. The authors also frame external, rule-based oversight as more reliable than self-critique for guiding agent behavior, and they argue that specification clarity—not raw task complexity—is the dominant lever on agent success.
Real-world applications the paper's framing points to:
- Computer-use and GUI automation agents that must click, drag, and manipulate interfaces at coordinate level rather than through high-level API calls.
- Web automation in the vein of WebArena and MacroBench, where models write or execute scripts against existing UI elements.
- Embodied or digital assistants asked to arrange objects or manipulate a digital environment from natural-language instructions.
- Safety-critical, human-in-the-loop deployments (the paper names industrial control and medical diagnosis) where each agent decision needs an auditable, rule-based justification.
Industry relevance: the results suggest text-only LLM agents can handle spatial planning without visual perception, that smaller models (GPT-4.1-mini at 90.8% perfect) remain competitive, and that external evaluators are a practical mechanism for improving agent output in production loops.
Future Directions
- Human performance baselines: the paper states it currently lacks these, so LLM scores cannot be contextualized against human ability.
- Beyond objective correctness: the rule-based system covers tool usage, coordinates, and coverage but cannot assess aesthetic quality of drawings.
- Adding vision: the text-only setting excludes visual feedback; the authors suggest screenshot-based feedback could improve iterative refinement.
- Long-horizon planning and tool state tracking: Very Hard tasks (60% perfect) exposed grid misalignment, miscounting, and failures to track which tool was selected, and the authors call for explicit state mechanisms for agents (they note that even with correct actions, models forgot to select required tools).
Target Audience
Researchers and practitioners working on LLM evaluation, agentic AI, GUI/web automation, and spatial or multimodal reasoning. It is also relevant to teams deploying autonomous agents in settings that demand auditable, reproducible scoring, and to benchmark designers looking for a template for transparent, externally supervised evaluation.
Authors’ abstract
As agentic AI systems increasingly operate autonomously, establishing trust through verifiable evaluation becomes critical. Yet existing benchmarks lack the transparency and auditability needed to assess whether agents behave reliably. We present DrawingBench, a verification framework for evaluating the trustworthiness of agentic LLMs through spatial reasoning tasks that require generating sequences of low-level GUI actions. Unlike opaque evaluations, DrawingBench provides transparent, rule-based assessment: 8 objective criteria enable reproducible scoring, while action-level inspection allows stakeholders to audit agent behavior. Our framework comprises 250 diverse prompts across 20 categories and 4 difficulty levels, deterministic evaluation metrics, and an external oversight mechanism through multi-turn feedback that enables human control over agent refinement. Evaluating four state-of-the-art LLMs (Claude-4 Sonnet, GPT-4.1, GPT-4.1-mini, Gemini-2.5 Flash) across 1,000 tests, we establish both capabilities and limitations: models achieved 92.8% perfect performance with structured external feedback driving significant improvements (average +3.2%, up to +32.8% for complex scenes), but systematic error patterns emerged in tool state management and long-horizon planning. Notably, specification clarity proved more important than task complexity -- models achieved 100% perfect performance when given explicit, verifiable criteria. These findings demonstrate that transparent evaluation frameworks can establish trust in agentic systems, with external oversight proving more reliable than self-correction for guiding agent behavior. Our open-source framework provides a template for trustworthy agent assessment. Code and data: https://github.com/hyunjun1121/DrawingBench