Most human-agent interaction today remains text-based. Natural language can impose cognitive overload, ambiguity, information chaos, and slow input for complex tasks; ephemeral generative UIs can present structured information and guide users toward task completion. We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces. Training the coder with reinforcement learning is challenging: verifiable rewards for interactive UI generation require costly execution, while LLM-as-a-Judge rewards are prone to reward hacking. We address the first challenge with Dynamic UX, a lightweight package for dynamic interaction and reward collection in a single sandbox, and the second with Reward Auditor, a meta-reward mechanism that monitors reward distributions and distills diagnostic patterns into a shared rubric and scoring specification. We introduce UI-TAU Bench, a benchmark for active human-agent interaction through generated UI code, built on 10 real-world domain databases constructed from public data sources and based on Tau-Bench tool-use settings, with Lite (300 tasks) and Full (1,000 tasks) splits. GenUI-Harness achieves an average Pass@3 gain of 4.48 percentage points over smolagents on Lite. Training with GenUI-Harness improves a 4B backbone from 9.33% to 58.00% Pass@3, outperforming larger frontier models such as Claude Opus 5 (46.67%). GenUI-Harness also remains robust on ambiguous and non-ambiguous queries. In a reviewer survey comparing communication channels, generated UIs reduce average dialogue rounds from 3.4 to 1.2. These results show that data-aware generative interfaces can support effective task completion and reduce dialogue rounds in evaluated database-backed workflows.
Figures & tables
Figure 1: Text-based and generated-UI interaction for the same flight-booking request. Left: text requires successive clarification of travel preferences. Right: the generated UI presents retrieved options and structured controls for direct review and clarification.
Figure 2: UI-Tau Bench construction pipeline. We collect and normalize the domain databases, construct and validate the query and action tools, and synthesize and validate the benchmark tasks.
Statistic
Lite
Full
Total instances
300
1,000
Moderate-ambiguity instances
161
523
High-ambiguity instances
139
477
Unique query tools
303
304
Unique action tools
162
331
Avg. tokens / user task
90.3
90.2
Table 1: Data statistics
Figure 3: GenUI -Harness links three stages: the Tool Agent retrieves task context, the GUI Coder Agent generates an interface to collect user input, and the Tool Agent uses that input to execute the final database action.
Figure 4: Reward Auditor adds an offline reward-revision loop to RL. Task outcomes on selected internal rollouts guide revisions to judging criteria and scoring rules.
GenUI -Harness
smolagents
mini-swe-agent
Pi
Model
Pass@3
Avg@3
Pass@3
Avg@3
Pass@3
Avg@3
Pass@3
Avg@3
Open-weight models
Qwen3.5-4B
9.33
3.22
10.67
3.90
11.67
4.11
13.67
5.00
Qwen3.6-27B
54.33
32.44
51.67
30.20
43.33
27.67
43.33
28.00
Qwen3.6-35B-A3B
47.33
26.89
43.67
25.90
36.33
20.11
38.00
21.78
gpt-oss-120b
47.67
28.22
30.00
14.10
50.00
28.22
40.33
23.11
Table 2: State EM Pass@3 and Avg@3 (%) on the 300-task Lite split across four harnesses (three sessions; temperature 0.7). GenUI -Harness uses generated UI; the other harnesses retain native text interaction. Bold marks each row’s best result; ties are bolded jointly.
Figure 7
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Public data source
Domain entities or workflows
License / terms
Flight booking
OpenFlights airport, airline, and route data ( OpenFlights, n.d. )
Airports, airlines, routes, equipment, and route-level connectivity
ODbL; component terms vary
Hotel booking
Hotel Booking Demand datasets ( António et al., 2019 )
Hotel bookings, arrival dates, room assignment, cancellation status, and special requests
CC BY 4.0
Bank account
Synthetic Mobile Money Transaction Dataset ( Azamuke, 2024 )
Account identifiers, transfers, transaction amounts, and balance transitions
Customers, sellers, products, orders, payments, delivery status, and reviews
CC BY-NC-SA 4.0
Movie ticket
MovieLens 25M ( Harper & Konstan, 2015 )
Pseudonymous users, movies, ratings, tags, and preference signals
GroupLens research-use terms
Appendix
Table 4: Public data sources for the 10 evaluation domains. The table summarizes the domain information provided by each source and its license or access terms.
Benchmark
Reported scale
Interaction setting
Evaluation target
Tau-Bench ( Yao et al., 2025 )
165 tasks; 2 domains
Agent converses with a simulated user and calls domain APIs
Final database state versus annotated goal
ToolSandbox ( Lu et al., 2025 )
1,032 scenarios; 34 tools
Agent converses with a simulated user and executes stateful tools
Intermediate and final milestones; forbidden events
AppWorld ( Trivedi et al., 2024 )
750 tasks; 9 apps; 457 APIs
Agent writes and executes programs to operate app APIs
State-based unit tests for task completion and unintended changes
WebLINX ( Lu et al., 2024 )
Approximately 2,300 demonstrations; over 150 websites
Agent follows dialogue instructions to navigate existing websites
Next-action prediction against expert demonstrations
UI-Tau Bench
300 Lite and 1,000 Full tasks; 10 domains
Agent generates an executable UI for user input, then acts through tools
Final database-state exact match after UI interaction and action execution
Appendix
Table 5: Comparison with related benchmarks, using the settings and scale reported in the cited papers. Tasks, scenarios, and demonstrations are distinct dataset units. Generated UI denotes front-end code produced for the user to operate, rather than code used by the agent to call APIs or navigate an existing interface.
Domain
Domain-specific tables
Columns
Records
Lite
Full
Flight booking
flights, bookings, airports
56
156
30
100
Hotel booking
hotels, room inventory, bookings
58
70
30
100
Bank account
accounts, transactions, loans
67
70
30
100
Restaurant reservation
restaurants, reservations, reviews
63
70
30
100
E-commerce
products, orders, product reviews
52
70
30
100
Movie ticket
movies, theaters, bookings
60
70
30
100
Appendix
Table 6: UI-Tau Bench domain statistics. Every domain provides a user-facing entity table, resource tables, and a transaction table. Columns and records count fields and rows across these tables; the table also reports the retained tasks in each split.
Split
Databases
Tasks
Tool Agent data
GUI-Coder data
SFT
25
12,376 / 254
32,409 / 669
12,376 / 254
RL
15
8,116 / 169
–
8,116 / 169
Appendix
Table 7: Training-domain split and data volume (train/validation). Each task contributes one GUI-Coder example or RL prompt; an SFT task can contribute multiple Tool Agent examples.
Setting
Tool Agent
GUI Coder Agent
Training / validation examples
32,409 / 669
12,376 / 254
Parameter update
Full fine-tuning
Full fine-tuning
Optimizer
AdamW
AdamW
Peak learning rate
1×10−5
1×10−5
Weight decay / gradient clip
0.05 / 1.0
0.05 / 1.0
Global batch size
32
32
Appendix
Table 8: Supervised fine-tuning configuration for the two role-specific agents.
Setting
Value
Initialization
GUI Coder Agent SFT checkpoint
Training / validation prompts
1,500 / 41
Optimizer
AdamW; learning rate 5×10−6 ; weight decay 0
Update budget
500 updates; 32 rollouts per update
Grouped sampling
8 rollouts per prompt; oversampling factor 2.0
Sequence limit
16,384 tokens
Appendix
Table 9: Shared GRPO configuration for the three independent GUI Coder Agent sessions.
Baseline
Δ Pass@3 (pp), 95% CI
Δ Avg@3 (pp), 95% CI
smolagents
+4.48 [2.37, 6.59]
+3.56 [2.15, 5.00]
mini-swe-agent
+5.81 [3.37, 8.26]
+3.00 [1.46, 4.58]
Pi
+7.15 [4.30, 9.96]
+4.04 [2.30, 5.81]
Appendix
Table 10: Macro-average paired differences across the nine shared backbones on Lite. Intervals are 95% task-paired bootstrap confidence intervals over 10,000 resamples.
Model
GenUI -Harness
smolagents
mini-swe-agent
Pi
Open-weight models
Qwen3.5-4B
4.11
3.80
4.70
5.30
Qwen3.6-27B
37.50
34.50
21.00
27.00
Qwen3.6-35B-A3B
35.30
27.10
24.50
21.90
gpt-oss-120b
32.20
14.50
30.00
25.50
DeepSeek V4 Flash
33.00
31.80
29.40
28.80
Appendix
Table 11: State EM Pass@1 (single-session Success Rate, %) on the Full split of UI-Tau Bench (1,000 tasks) across four agent harnesses in their native interaction modes. Bold marks the best result in each row.
C
System
Renders/min ↑
Median (ms) ↓
P95 (ms) ↓
Peak RSS (MB) ↓
Chromium processes ↓
1
Baseline
25.48
2,237.5
2,349.1
894.7
7
Dynamic UX
26.47 ( ↑ 3.89%)
2,226.5 ( ↓ 0.49%)
2,328.8 ( ↓ 0.86%)
876.1 ( ↓ 2.08%)
6 ( ↓ 14.29%)
4
Baseline
94.30
2,471.8
2,863.4
2,306.8
24
Dynamic UX
101.23 ( ↑ 7.35%)
2,271.2 ( ↓ 8.12%)
2,382.5 ( ↓ 16.79%)
1,556.3 ( ↓ 32.53%)
10 ( ↓ 58.33%)
8
Baseline
144.18
3,188.8
4,076.9
4,388.7
48
Dynamic UX
189.03 ( ↑ 31.11%)
2,410.1 ( ↓ 24.42%)
2,548.1 ( ↓ 37.50%)
2,525.0 ( ↓ 42.47%)
15 ( ↓ 68.75%)
Appendix
Table 12: Matched-concurrency rendering benchmark. Baseline launches one browser per rollout; Dynamic UX shares one browser across isolated contexts. Bold marks the better value; green arrows show relative change from Baseline.
Diagnosed problem
Reward revision
Intended effect
Plausible interfaces receive high scores despite interaction or payload defects
Revise judging criteria for actionable controls, submission wiring, payload semantics, and task fidelity
Make task-blocking defects explicit in the reward
Different failure stages collapse to the same zero score
Assign stage-aware low-reward bands
Distinguish progress within unsuccessful rollouts
Generated code uses unavailable icon exports
Apply a deterministic unsupported-icon gate
Prevent the quality score from overriding the detected defect
Label forms or associations impede control access
Apply source-based accessibility penalties
Include interaction-relevant defects in score assignment
Appendix
Table 13: Problems targeted by meta-reward-guided reward revision. Revisions combine judging criteria with deterministic scoring rules.
Detected defect
Maximum judge score
Empty/error interface or no actionable control
All four dimensions =0
Missing or unreachable submitData
FE, TA ≤1
Missing required controls
FC ≤4,3,2 when one, two, or at least half are missing
Wrong, constant, or omitted submitted values
FE, TA ≤3
Broken control wiring
IC ≤3 ; TA ≤3 when task-blocking
Harmful invented facts
TA ≤2
Appendix
Table 14: Explicit judge caps in the Audit Round 1 revision. FE, FC, IC, and TA denote functional equivalence, feature completeness, implementation correctness, and task alignment.
Configuration
No valid interact.
Wrong tool
Wrong args
Exec. rejected
Success
DeepSeek V4 Flash
124
0
71
7
98
Claude Sonnet 4.6
87
0
102
10
101
Claude Opus 5
82
0
124
1
93
GenUI -4B (Ours)
123
0
59
10
108
Appendix
Table 15: Single-session outcome partition on the 300 Lite tasks under GenUI -Harness. Columns are mutually exclusive and sum to 300 for every configuration.
Mechanism
Trace evidence
Render failure
Compiler or runtime errors prevent an interactive page from loading.
Inaccessible control
A control’s rendered accessible name or label binding does not support the attempted browser interaction.
Insufficient input controls
The available controls cannot express the required selection or value.
Incorrect submitted value
The payload diverges from the task requirement or displayed choice, with no control to correct it.
Appendix
Table 16: Observed failure mechanisms and the evidence used to identify them.
Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an executable artifact as user requests evolve. We introduce EvoGenUI-Bench, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state. We execute generated artifacts in a browser and evaluate them using screenshots, source and DOM evidence, actor traces, and runtime logs. Beyond turn-level and episode-level success, we measure cross-turn retention with Adjacent Pass Retention. Across eight models, even the strongest achieves 74.9% Turn Pass while completing only 37.3% of five-turn episodes; APR further falls to 52.4% on tool-grounded tasks. Diagnostic analysis shows that presentation failures center on information architecture, interaction failures on derived-state propagation and affordance binding, and tool-grounded failures additionally involve external-state grounding and requirement decomposition. These results reframe generative UI evaluation from judging isolated outputs to testing whether interface behavior, derived state, external state, and assistant claims remain synchronized as the artifact evolves.
Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson r from 0.716 to 0.922, and a prompt-ensemble control recovers only about a third of this gap, isolating genuine persona and evidence grounding as the dominant source of improvement. Beyond this fidelity gain, retaining each panelist's individual rating further reveals that user subgroups agree on overall model rankings yet diverge sharply on specific rating dimensions, a structural disagreement a single homogeneous judge would systematically erase. The codes are available at https://github.com/Wuzheng02/ESPP.
Zheng Wu, Yibo Luo, Pu Zhang +2
School of Computer Science, Shanghai Jiao Tong University · 2ByteDance Inc
GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.