Research
When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
Overview Research area: Human-agent interaction, generative user interfaces, tool-using LLM agents, and reinforcement learning for code generation. Technical level: Advanced. Scope: The paper proposes
- arXiv
- 2610.11123
- Published
- 2026-10-08
- Authors
- Xiaolong Li, Xiaohan Xu, Jinyang Li, Xinnuo Xu, Ge Qu, Nan Huo, Jack Williams, Reynold Cheng
AI summary
Overview
Research area: Human-agent interaction, generative user interfaces, tool-using LLM agents, and reinforcement learning for code generation.
Technical level: Advanced.
Scope: The paper proposes GenUI-Harness, a multi-agent system in which an agent that retrieves database records and an agent that writes React/TypeScript interface code collaborate to replace multi-turn text dialogue with an ephemeral, task-specific UI, and which is evaluated on a new benchmark called UI-Tau Bench.
What This Paper Is About
Most human-agent interaction is still text-based, and the authors argue that natural language is a poor fit for complex, goal-directed tasks because it causes cognitive overload, ambiguity, scattered information, and slow input. Their goal is to let an agent generate a small executable interface on the fly, grounded in records it has already retrieved from a database, so the user can review real options and submit missing choices in one structured exchange instead of several rounds of dialogue. The generated submission is then used by the agent to execute a database action whose resulting state is checked against a target state.
Key Contributions
-
The task of data-aware generative UI for stateful service tasks. The agent must construct an executable interface from a user request and database evidence, collect the user input needed, and complete the requested action through tool use. Each instance contains a user request, an initial database state, query and action tools, and domain rules, following the tool-use setting of Tau-Bench.
-
UI-Tau Bench. A benchmark for active human-agent interaction through generated front-end code, built on 10 real-world domain databases constructed from public data sources, with Lite (300 tasks) and Full (1,000 tasks) splits. It reports Pass@3 and Avg@3 state exact-match metrics over three evaluation sessions per task.
-
GenUI-Harness. A harness combining tool grounding, role-specialized UI generation, Dynamic UX, and Reward Auditor. A Tool Agent handles pre-agent exploration and post-agent execution; a GUI Coder Agent emits a single Material UI component in React and TypeScript that presents retrieved candidates and collects unresolved arguments.
-
Two training enablers. Dynamic UX provides parallel, session-isolated execution and reward collection within a single sandbox; Reward Auditor is a meta-reward mechanism that uses actual task-execution outcomes to diagnose and automatically revise the training reward.
Main Findings
-
Harness-level gains on Lite. Across nine backbones, GenUI-Harness leads Pass@3 on seven of nine backbones and improves Pass@3 over smolagents by 4.48 percentage points in macro-average (task-paired bootstrap 95% CI: [2.37, 6.59]). It attains higher Pass@3 than smolagents, mini-swe-agent, and Pi on 8, 7, and 8 of the 9 shared backbones respectively.
-
Compact model beats larger frontier models after training. SFT increases Avg@3 from 3.22% to 27.67%. With the final audited reward, GenUI-4B reaches 58.00% Pass@3 and 36.00% Avg@3, up from a 9.33% Pass@3 baseline for the 4B backbone, and above Claude Opus 5 at 46.67% Pass@3 and 31.56% Avg@3.
-
Fewer agent turns. Among successful episodes, GenUI-Harness achieves higher Avg@3 with 36–43% fewer agent turns, placing it alone on the favorable lower-right frontier of the Avg@3 versus agent-turns plot.
-
Generated UIs cut dialogue rounds. In a reviewer survey, generated UIs reduce the average number of dialogue rounds from 3.4 to 1.2. UI latency is judged acceptable on 85% of tasks by a majority of UI reviewers.
-
GUI-Coder transfers across reasoning models. A code-stage intervention holding pre-agent output fixed and replacing only code generation improves Lite Avg@3 for all five tested models: Qwen3.6-35B-A3B 26.89 to 29.33 (+2.44), DeepSeek V4 Flash 30.22 to 38.67 (+8.45), GPT-5.4 25.22 to 30.00 (+4.78), Claude Sonnet 4.6 35.67 to 42.33 (+6.66), and Claude Opus 5 31.56 to 40.33 (+8.77).
-
Robust on unambiguous requests. Non-ambiguous Lite variants were created by revealing the user-facing arguments masked in the original requests while keeping tools, databases, user simulator, and evaluation protocol fixed; GenUI-4B remains strongest in both settings.
-
Failure modes are interface operability and argument fidelity. Across four models on the 300 Lite tasks, failures are dominated by unsuccessful UI interaction (52.0%) and mismatched action arguments (44.5%). Examples include a single-select control that cannot express a required multi-selection, and an immutable cabin-class field that submits the wrong value.
-
Dynamic UX is more efficient than fresh-process rendering. Benchmarked on 1,100 pre-verified components across five concurrency levels, it increases throughput by 18.2% while reducing median latency by 15.3%, peak RSS by 35.2%, and Chromium process count by 59.4%. A 1,000-session marker test completes without timeout, failure, or content leakage.
-
Benchmark statistics. Lite/Full splits contain 300/1,000 total instances, 161/523 moderate-ambiguity instances, 139/477 high-ambiguity instances, 303/304 unique query tools, 162/331 unique action tools, 90.3/90.2 average tokens per user task, 4.9/4.8 average masked arguments per action, and inter-annotator agreement of 94.1%/94.3%.
Methodology in Plain English
The system splits the work between two specialists. A Tool Agent first explores the database with read-only tool calls (up to 5 turns in a ReAct loop) to find relevant records, candidate options, and the user profile. That evidence, along with the query and the exploration history, is handed to a GUI Coder Agent, which decides which action arguments the records already settle, which candidates are available for the rest, and which only the user can decide. It then writes one React/TypeScript component with controls suited to those decisions, showing resolved arguments as context or prefilled values so the user only attends to what is unresolved. The user simulator operates the rendered interface according to its assigned goal, and its submission returns to the Tool Agent as a JSON object; the Tool Agent can use up to 5 further query turns, then calls the action tool. The evaluator resets an isolated database copy and checks whether the final database state exactly matches the target.
Training is done separately from the benchmark. Forty databases were synthesized with Claude Sonnet 4.6, and DeepSeek V4 Flash generated end-to-end trajectories inside the harness; only trajectories whose final database state matched the target and whose interfaces rendered with valid ARIA snapshots were kept. Trajectories from 25 databases were used for supervised fine-tuning of the Tool Agent and GUI Coder Agent separately, both starting from Qwen3.5-4B, producing GenUI-4B (SFT). The remaining 15 databases define the RL environments. Holding the Tool Agent fixed at its SFT checkpoint, only the GUI Coder Agent was trained with GRPO.
Two problems motivated the reward design. First, collecting execution-based rewards requires rendering generated code at scale with isolated sessions, which Dynamic UX solves by sharing one rendering runtime and browser process while giving each rollout an isolated browser context. Second, rendering alone does not show whether an interface exposes the relevant ambiguities or enables successful downstream action, and an LLM-as-a-Judge can assign high scores to interfaces that fail in use. Reward Auditor therefore uses execution outcomes to diagnose mismatches: the objective label is positive only when execution succeeds and the database state matches the target, compared against a reward prediction thresholded at 0.5. The auditor, using Claude Opus 5, examines reward distributions, judge rationales, and complete trajectories, proposes revisions to a shared rubric and scoring specification, and replays candidate rewards on stored rollouts with generated code held fixed so score changes are attributable to the reward revision. Two rounds of revision produced the final training reward. The initial reward was a weighted judge score (0.35 functional equivalence, 0.35 feature completeness, 0.15 implementation correctness, 0.15 task alignment, each in [0, 5]) gated by validity checks.
Why This Matters
Impact on research. The work reframes the interface channel itself as a learnable component of an agent system rather than a fixed text box, and it connects retrieved records, executable interaction, and terminal state evaluation within one task. It also contributes a training methodology for UI generation where verifiable rewards are expensive: execution-gated rollouts plus a meta-reward auditing loop that revises the reward from observed task outcomes rather than trusting a judge proxy.
Real-world applications.
- Flight booking: the motivating example consolidates retrieved flight candidates, pre-grounded constraints, and missing slots into a single canvas, replacing successive clarification turns.
- Hotel booking, e-commerce, and retail: these are among the 10 real-world domain databases in UI-Tau Bench, all constructed from public data sources including OpenFlights, António et al. (2019), Olist (2018), and Chen (2012).
- Customer-facing service agents that must update a backend record (change a booking, place an order) while satisfying explicit domain rules and eligibility conditions.
- Internal tools where an agent explains retrieved database state to a user and captures a structured decision, avoiding exposure of schemas or internal identifiers.
Industry relevance. The results show a 4B backbone trained in the harness reaching 58.00% Pass@3, exceeding Claude Opus 5 at 46.67% on the same Lite split, and GUI-Coder improves five other reasoning models when swapped in. That suggests a specialised, compact interface-generation model is a deployable component, and the Dynamic UX measurements (18.2% higher throughput, 59.4% fewer Chromium processes) speak to the infrastructure cost of training and serving such components.
Future Directions
-
Validate with real users rather than a simulator. The paper uses a fixed, capable-user simulator and states it does not claim to reproduce the full distribution of human interaction behavior; the reviewer survey covers 100 tasks with three reviewers per channel, so broader human studies are an open question.
-
Fix the interface-operability bottleneck. Unsuccessful UI interaction accounts for 52.0% of failures across four models on the 300 Lite tasks, with concrete defects such as single-select controls that cannot express required multi-selections.
-
Close the argument-fidelity gap. Mismatched action arguments account for 44.5% of failures, which the authors say requires examining both submissions and resulting database changes.
-
Extend evaluation beyond the reported Lite results. Full-split (1,000-task) results are deferred to Appendix D.6 rather than reported in the main text, and the reported comparisons cover nine backbones on Lite.
Target Audience
Researchers and practitioners working on LLM agents, tool use, and stateful task benchmarks; reinforcement learning and reward-design researchers interested in execution-gated rewards and reward hacking in judge-based feedback; and HCI and interface-generation researchers studying how generated or task-specific interfaces mediate human-agent interaction.
Authors’ abstract
Most human-agent interaction today remains text-based. Natural language can impose cognitive overload, ambiguity, information chaos, and slow input for complex tasks; ephemeral generative UIs can present structured information and guide users toward task completion. We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces. Training the coder with reinforcement learning is challenging: verifiable rewards for interactive UI generation require costly execution, while LLM-as-a-Judge rewards are prone to reward hacking. We address the first challenge with Dynamic UX, a lightweight package for dynamic interaction and reward collection in a single sandbox, and the second with Reward Auditor, a meta-reward mechanism that monitors reward distributions and distills diagnostic patterns into a shared rubric and scoring specification. We introduce UI-TAU Bench, a benchmark for active human-agent interaction through generated UI code, built on 10 real-world domain databases constructed from public data sources and based on Tau-Bench tool-use settings, with Lite (300 tasks) and Full (1,000 tasks) splits. GenUI-Harness achieves an average Pass@3 gain of 4.48 percentage points over smolagents on Lite. Training with GenUI-Harness improves a 4B backbone from 9.33% to 58.00% Pass@3, outperforming larger frontier models such as Claude Opus 5 (46.67%). GenUI-Harness also remains robust on ambiguous and non-ambiguous queries. In a reviewer survey comparing communication channels, generated UIs reduce average dialogue rounds from 3.4 to 1.2. These results show that data-aware generative interfaces can support effective task completion and reduce dialogue rounds in evaluated database-backed workflows.