When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
Organizations: The University of Hong Kong · The BIRD Team · Microsoft Research
Abstract
Most human-agent interaction today remains text-based. Natural language can impose cognitive overload, ambiguity, information chaos, and slow input for complex tasks; ephemeral generative UIs can present structured information and guide users toward task completion. We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces. Training the coder with reinforcement learning is challenging: verifiable rewards for interactive UI generation require costly execution, while LLM-as-a-Judge rewards are prone to reward hacking. We address the first challenge with Dynamic UX, a lightweight package for dynamic interaction and reward collection in a single sandbox, and the second with Reward Auditor, a meta-reward mechanism that monitors reward distributions and distills diagnostic patterns into a shared rubric and scoring specification. We introduce UI-TAU Bench, a benchmark for active human-agent interaction through generated UI code, built on 10 real-world domain databases constructed from public data sources and based on Tau-Bench tool-use settings, with Lite (300 tasks) and Full (1,000 tasks) splits. GenUI-Harness achieves an average Pass@3 gain of 4.48 percentage points over smolagents on Lite. Training with GenUI-Harness improves a 4B backbone from 9.33% to 58.00% Pass@3, outperforming larger frontier models such as Claude Opus 5 (46.67%). GenUI-Harness also remains robust on ambiguous and non-ambiguous queries. In a reviewer survey comparing communication channels, generated UIs reduce average dialogue rounds from 3.4 to 1.2. These results show that data-aware generative interfaces can support effective task completion and reduce dialogue rounds in evaluated database-backed workflows.
Figures & tables
| Statistic | Lite | Full |
| Total instances | 300 | 1,000 |
| Moderate-ambiguity instances | 161 | 523 |
| High-ambiguity instances | 139 | 477 |
| Unique query tools | 303 | 304 |
| Unique action tools | 162 | 331 |
| Avg. tokens / user task | 90.3 | 90.2 |
| GenUI -Harness | smolagents | mini-swe-agent | Pi | |||||
| Model | Pass@3 | Avg@3 | Pass@3 | Avg@3 | Pass@3 | Avg@3 | Pass@3 | Avg@3 |
| Open-weight models | ||||||||
| Qwen3.5-4B | 9.33 | 3.22 | 10.67 | 3.90 | 11.67 | 4.11 | 13.67 | 5.00 |
| Qwen3.6-27B | 54.33 | 32.44 | 51.67 | 30.20 | 43.33 | 27.67 | 43.33 | 28.00 |
| Qwen3.6-35B-A3B | 47.33 | 26.89 | 43.67 | 25.90 | 36.33 | 20.11 | 38.00 | 21.78 |
| gpt-oss-120b | 47.67 | 28.22 | 30.00 | 14.10 | 50.00 | 28.22 | 40.33 | 23.11 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Public data source | Domain entities or workflows | License / terms |
|---|---|---|---|
| Flight booking | OpenFlights airport, airline, and route data ( OpenFlights, n.d. ) | Airports, airlines, routes, equipment, and route-level connectivity | ODbL; component terms vary |
| Hotel booking | Hotel Booking Demand datasets ( António et al., 2019 ) | Hotel bookings, arrival dates, room assignment, cancellation status, and special requests | CC BY 4.0 |
| Bank account | Synthetic Mobile Money Transaction Dataset ( Azamuke, 2024 ) | Account identifiers, transfers, transaction amounts, and balance transitions | CC BY 4.0 |
| Restaurant reservation | UCI Restaurant & Consumer Data ( Medellín & Serna, 2011 ) | User profiles and preferences, restaurants, accepted payments, hours, and ratings | CC BY 4.0 |
| E-commerce | Olist Brazilian E-Commerce Dataset ( Olist, 2018 ) | Customers, sellers, products, orders, payments, delivery status, and reviews | CC BY-NC-SA 4.0 |
| Movie ticket | MovieLens 25M ( Harper & Konstan, 2015 ) | Pseudonymous users, movies, ratings, tags, and preference signals | GroupLens research-use terms |
| Benchmark | Reported scale | Interaction setting | Evaluation target |
|---|---|---|---|
| Tau-Bench ( Yao et al., 2025 ) | 165 tasks; 2 domains | Agent converses with a simulated user and calls domain APIs | Final database state versus annotated goal |
| ToolSandbox ( Lu et al., 2025 ) | 1,032 scenarios; 34 tools | Agent converses with a simulated user and executes stateful tools | Intermediate and final milestones; forbidden events |
| AppWorld ( Trivedi et al., 2024 ) | 750 tasks; 9 apps; 457 APIs | Agent writes and executes programs to operate app APIs | State-based unit tests for task completion and unintended changes |
| WebLINX ( Lu et al., 2024 ) | Approximately 2,300 demonstrations; over 150 websites | Agent follows dialogue instructions to navigate existing websites | Next-action prediction against expert demonstrations |
| UI-Tau Bench | 300 Lite and 1,000 Full tasks; 10 domains | Agent generates an executable UI for user input, then acts through tools | Final database-state exact match after UI interaction and action execution |
| Domain | Domain-specific tables | Columns | Records | Lite | Full |
|---|---|---|---|---|---|
| Flight booking | flights, bookings, airports | 56 | 156 | 30 | 100 |
| Hotel booking | hotels, room inventory, bookings | 58 | 70 | 30 | 100 |
| Bank account | accounts, transactions, loans | 67 | 70 | 30 | 100 |
| Restaurant reservation | restaurants, reservations, reviews | 63 | 70 | 30 | 100 |
| E-commerce | products, orders, product reviews | 52 | 70 | 30 | 100 |
| Movie ticket | movies, theaters, bookings | 60 | 70 | 30 | 100 |
| Split | Databases | Tasks | Tool Agent data | GUI-Coder data |
|---|---|---|---|---|
| SFT | 25 | 12,376 / 254 | 32,409 / 669 | 12,376 / 254 |
| RL | 15 | 8,116 / 169 | – | 8,116 / 169 |
| Setting | Tool Agent | GUI Coder Agent |
|---|---|---|
| Training / validation examples | 32,409 / 669 | 12,376 / 254 |
| Parameter update | Full fine-tuning | Full fine-tuning |
| Optimizer | AdamW | AdamW |
| Peak learning rate | ||
| Weight decay / gradient clip | 0.05 / 1.0 | 0.05 / 1.0 |
| Global batch size | 32 | 32 |
| Setting | Value |
|---|---|
| Initialization | GUI Coder Agent SFT checkpoint |
| Training / validation prompts | 1,500 / 41 |
| Optimizer | AdamW; learning rate ; weight decay 0 |
| Update budget | 500 updates; 32 rollouts per update |
| Grouped sampling | 8 rollouts per prompt; oversampling factor 2.0 |
| Sequence limit | 16,384 tokens |
| Baseline | Pass@3 (pp), 95% CI | Avg@3 (pp), 95% CI |
|---|---|---|
| smolagents | +4.48 [2.37, 6.59] | +3.56 [2.15, 5.00] |
| mini-swe-agent | +5.81 [3.37, 8.26] | +3.00 [1.46, 4.58] |
| Pi | +7.15 [4.30, 9.96] | +4.04 [2.30, 5.81] |
| Model | GenUI -Harness | smolagents | mini-swe-agent | Pi |
|---|---|---|---|---|
| Open-weight models | ||||
| Qwen3.5-4B | 4.11 | 3.80 | 4.70 | 5.30 |
| Qwen3.6-27B | 37.50 | 34.50 | 21.00 | 27.00 |
| Qwen3.6-35B-A3B | 35.30 | 27.10 | 24.50 | 21.90 |
| gpt-oss-120b | 32.20 | 14.50 | 30.00 | 25.50 |
| DeepSeek V4 Flash | 33.00 | 31.80 | 29.40 | 28.80 |
| C | System | Renders/min | Median (ms) | P95 (ms) | Peak RSS (MB) | Chromium processes |
|---|---|---|---|---|---|---|
| 1 | Baseline | 25.48 | 2,237.5 | 2,349.1 | 894.7 | 7 |
| Dynamic UX | 26.47 ( 3.89%) | 2,226.5 ( 0.49%) | 2,328.8 ( 0.86%) | 876.1 ( 2.08%) | 6 ( 14.29%) | |
| 4 | Baseline | 94.30 | 2,471.8 | 2,863.4 | 2,306.8 | 24 |
| Dynamic UX | 101.23 ( 7.35%) | 2,271.2 ( 8.12%) | 2,382.5 ( 16.79%) | 1,556.3 ( 32.53%) | 10 ( 58.33%) | |
| 8 | Baseline | 144.18 | 3,188.8 | 4,076.9 | 4,388.7 | 48 |
| Dynamic UX | 189.03 ( 31.11%) | 2,410.1 ( 24.42%) | 2,548.1 ( 37.50%) | 2,525.0 ( 42.47%) | 15 ( 68.75%) |
| Diagnosed problem | Reward revision | Intended effect |
|---|---|---|
| Plausible interfaces receive high scores despite interaction or payload defects | Revise judging criteria for actionable controls, submission wiring, payload semantics, and task fidelity | Make task-blocking defects explicit in the reward |
| Different failure stages collapse to the same zero score | Assign stage-aware low-reward bands | Distinguish progress within unsuccessful rollouts |
| Generated code uses unavailable icon exports | Apply a deterministic unsupported-icon gate | Prevent the quality score from overriding the detected defect |
| Label forms or associations impede control access | Apply source-based accessibility penalties | Include interaction-relevant defects in score assignment |
| Detected defect | Maximum judge score |
|---|---|
| Empty/error interface or no actionable control | All four dimensions |
| Missing or unreachable submitData | FE, TA |
| Missing required controls | FC when one, two, or at least half are missing |
| Wrong, constant, or omitted submitted values | FE, TA |
| Broken control wiring | IC ; TA when task-blocking |
| Harmful invented facts | TA |
| Configuration | No valid interact. | Wrong tool | Wrong args | Exec. rejected | Success |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | 124 | 0 | 71 | 7 | 98 |
| Claude Sonnet 4.6 | 87 | 0 | 102 | 10 | 101 |
| Claude Opus 5 | 82 | 0 | 124 | 1 | 93 |
| GenUI -4B (Ours) | 123 | 0 | 59 | 10 | 108 |
| Mechanism | Trace evidence |
|---|---|
| Render failure | Compiler or runtime errors prevent an interactive page from loading. |
| Inaccessible control | A control’s rendered accessible name or label binding does not support the attempted browser interaction. |
| Insufficient input controls | The available controls cannot express the required selection or value. |
| Incorrect submitted value | The payload diverges from the task requirement or displayed choice, with no control to correct it. |