Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench. However, binary preference prediction provides limited insight, while relying on a single user-agnostic reward model overlooks the inherent heterogeneity of users, whose expectations can differ substantially. In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains. Each instance is paired with a FACTORS user profile that characterizes the user through seven interpretable behavioral facets, differentiating user groups. To provide richer evaluation insights, we introduce a dual-perspective paradigm that combines a personalized User Reward Model (URM) for third-person judgment with Sim4Eval, a user simulator that enables multi-turn interactions and provides first-person evaluation across four cognitive state dimensions. To assess the reliability of these based evaluators, we further introduce two meta-benchmarks, URMBench and USimBench, that evaluate how faithfully they reproduce real human preferences and behaviors. Extensive experiments reveal seven key findings that highlight the importance of user modeling and multi-perspective evaluation, offering a fresh perspective on user-centric benchmarking and motivating personalized model optimization.
Figures & tables
Figure 1: Illustration of the seven-dimensional FACTORS profile schema for characterizing heterogeneous users in human-AI interactions.
Figure 2: Overview of UXBench Pro with three core proposed methodologies.
Dimension
Signal
Categories
Feedback ( F )
Likes/dislikes, 30d
Silent, critical, lenient, expressive, light
Age ( A )
Age interval
Adolescent, young adult, adult, senior, unknown
Consumption ( C )
Device tier
Low, medium, high, flagship, unknown
Technical Usage ( T )
Prompt volume + usage
Light, regular, heavy
Role ( O )
Dialogue-mined life stage
Student, working-age, retired, parenting, unknown
Region ( R )
Geographic location
Metro, town, rural, unknown
Table 1: FACTORS user profile schema, with each dimension discretized into interpretable categories plus an unknown state for completeness.
Figure 3: Distribution of the 1,000 test cases across each FACTORS facet.
Category
Code
Scenario
N
Work
W1
Office writing (documents, emails, reports)
107
W2
Coding (programming, debugging)
17
W3
Data analysis (analytics, decision support)
110
W4
Professional lookup (legal, financial, medical)
39
W5
Translation (business translation)
18
Life
L1
Study assistance (learning, homework)
104
Table 2: Task taxonomy and per-scenario instance counts in the released set.
Statistic
Value
Test instances
1,000
Task scenarios
12
Domains (coarse / fine)
15 / 82
Work / life split
29.1% / 70.9%
Normalized task entropy
0.953
Instances with a preceding dialogue context
63.2%
Table 3: Summary statistics of UXBench Pro.
Figure 4: Composition of the experience signals: (a) failure dimension, (b) user actions, (c) preceding dialogue turns.
URM (Pairwise GRM)
User state
Trajectory outcome
Assistant
Turn BT ↑
Traj. BT ↑
S1 ↑
S2 ↑
S3 ↑
S4 ↑
Res. ↑
Aban. ↓
Qwen3.7-Max
+0.170
+0.096
3.18
3.18
3.75
3.39
37.5
1.6
DeepSeek-v4-pro
+0.047
+0.053
3.24
3.22
3.84
3.44
41.4
1.9
Hunyuan-3
+0.028
−0.027
3.19
3.25
3.87
3.49
38.7
2.2
GLM-5.2
+0.010
−0.012
3.15
3.20
3.79
3.39
34.5
2.1
Claude-Opus-4.8
−0.014
−0.006
3.09
3.13
3.80
3.39
37.4
4.7
Table 4: UXBench Pro results for seven assistants. Turn BT and Traj. BT are Bradley–Terry strengths derived from pairwise URM judgments. Res. and Aban. denote resolution and abandonment rates. Separability is the fraction of assistant pairs with non-overlapping 95% bootstrap confidence intervals; ρ denotes Spearman correlation.
Formulation
Training signal
S1 ↑
S2 ↑
S3 ↑
S4 ↑
Pooled ↑
Pointwise GRM
Pointwise feedback labels
0.5500
0.5300
0.4650
0.5650
0.5275
General RM
Mixed, general purpose
0.6700
0.5300
0.4850
0.6100
0.5737
BTRM
Regeneration preference
0.6900
0.6100
0.4900
0.5900
0.5950
BTRM
Online dual-answer choice
0.6850
0.6050
0.5400
0.6250
0.6138
Pairwise GRM ✓
Regeneration preference
0.8650
0.8050
0.6200
0.7850
0.7688
Table 5: PairAcc of five trained reward models on the 800-instance URM-Bench under the population condition. Chance is 0.5, exact ties count as errors, and ✓ marks the model selected as the URM for UXBench Pro.
Low effort
High effort
System
Base
+Profile
Base
+Profile
DeepSeek-v4-pro
0.5875
0.5938
0.6175
0.6400
Hunyuan-3
0.6475
0.6625
0.6150
0.6212
GPT-5.5
0.6375
0.6462
0.5863
0.6175
Gemini-3.1-pro
0.6025
0.6225
0.5913
0.5887
Table 7: Thinking-effort ablation on four judges. Values are pooled PairAcc under low and high thinking effort, with and without profile conditioning. The best result for each model is bolded.
Base model
n
S1 ↓
S2 ↓
S3 ↓
S4 ↓
All four ↓
Generic scale
Deepseek-v4-Pro
117
0.248
0.504
0.368
0.376
0.120
Deepseek-v4-Flash
147
0.197
0.327
0.367
0.286
0.082
GPT-5.2
150
0.307
0.573
0.553
0.540
0.260
Hunyuan-3
150
0.833
0.947
0.987
0.953
0.820
Anchored variant
Table 9: State persistence in multi-turn RAG user simulation. Each value is the fraction of sessions in which a state coordinate, or all four jointly, remains unchanged after the first turn. n denotes the number of sessions with at least two scored turns.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Code
Scenario
Assignment signal
WK-a
Office writing
Workplace, business, or professional writing intent
WK-b
Programming
Code intent or programming meta-intent
WK-c
Analysis and decision
Analysis or decision intent; finance for working-role users
WK-d
Professional retrieval
Legal industry
WK-e
Business translation
Translation intent
LF-a
Learning assistance
Mathematics intent, education, or scholastic knowledge
Appendix
Table 10: Deterministic trace-level scenario assignment in UXBench Pro.
As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. We present \textbf{UXBench}, the first user-centric benchmark grounded in real user feedback signals for evaluating preference alignment and dialogue generation. The benchmark consists of three interconnected tasks, UX Judge, UX Eval, and UX Recovery, with 7,400 test instances extracted from over 70K interaction logs of a mainstream Chinese AI assistant. The dataset closely reflects real user distributions, covering 8 scenarios, 83 domains, and diverse failure patterns that pose severe challenges. Extensive experiments on 26 frontier language models provide novel insights into how well models perceive user experience and how improvements in model capability contribute to better dialogue engagement. Through analyses of model behavior and performance gaps, we document six important findings, demonstrating that user feedback prediction is a learnable capability and revealing different aspects that influence user experience. UXBench establishes a new evaluation landscape and calls for greater attention to tailored UX optimization, contributing toward a user-centric scaling law for the development of successful AI assistants. The full project is released at https://github.com/mengze-hong/UXBench.
Mengze Hong, Xia Zeng, Zeyang Lei +13
Hong Kong Polytechnic University · Yuanbao Team, Tencent
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.
Yuhang Zhu, Mingxuan Du, Benfeng Xu +3
University of Science and Technology of China · MetaStone · MetaStone Technology, Beijing, China