Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench. However, binary preference prediction provides limited insight, while relying on a single user-agnostic reward model overlooks the inherent heterogeneity of users, whose expectations can differ substantially. In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains. Each instance is paired with a FACTORS user profile that characterizes the user through seven interpretable behavioral facets, differentiating user groups. To provide richer evaluation insights, we introduce a dual-perspective paradigm that combines a personalized User Reward Model (URM) for third-person judgment with Sim4Eval, a user simulator that enables multi-turn interactions and provides first-person evaluation across four cognitive state dimensions. To assess the reliability of these based evaluators, we further introduce two meta-benchmarks, URMBench and USimBench, that evaluate how faithfully they reproduce real human preferences and behaviors. Extensive experiments reveal seven key findings that highlight the importance of user modeling and multi-perspective evaluation, offering a fresh perspective on user-centric benchmarking and motivating personalized model optimization.
Figures & tables
Figure 1: Illustration of the seven-dimensional FACTORS profile schema for characterizing heterogeneous users in human-AI interactions.
Figure 2: Overview of UXBench Pro with three core proposed methodologies.
Dimension
Signal
Categories
Feedback ( F )
Likes/dislikes, 30d
Silent, critical, lenient, expressive, light
Age ( A )
Age interval
Adolescent, young adult, adult, senior, unknown
Consumption ( C )
Device tier
Low, medium, high, flagship, unknown
Technical Usage ( T )
Prompt volume + usage
Light, regular, heavy
Role ( O )
Dialogue-mined life stage
Student, working-age, retired, parenting, unknown
Region ( R )
Geographic location
Metro, town, rural, unknown
Table 1: FACTORS user profile schema, with each dimension discretized into interpretable categories plus an unknown state for completeness.
Figure 3: Distribution of the 1,000 test cases across each FACTORS facet.
Category
Code
Scenario
N
Work
W1
Office writing (documents, emails, reports)
107
W2
Coding (programming, debugging)
17
W3
Data analysis (analytics, decision support)
110
W4
Professional lookup (legal, financial, medical)
39
W5
Translation (business translation)
18
Life
L1
Study assistance (learning, homework)
104
Table 2: Task taxonomy and per-scenario instance counts in the released set.
Statistic
Value
Test instances
1,000
Task scenarios
12
Domains (coarse / fine)
15 / 82
Work / life split
29.1% / 70.9%
Normalized task entropy
0.953
Instances with a preceding dialogue context
63.2%
Table 3: Summary statistics of UXBench Pro.
Figure 4: Composition of the experience signals: (a) failure dimension, (b) user actions, (c) preceding dialogue turns.
URM (Pairwise GRM)
User state
Trajectory outcome
Assistant
Turn BT ↑
Traj. BT ↑
S1 ↑
S2 ↑
S3 ↑
S4 ↑
Res. ↑
Aban. ↓
Qwen3.7-Max
+0.170
+0.096
3.18
3.18
3.75
3.39
37.5
1.6
DeepSeek-v4-pro
+0.047
+0.053
3.24
3.22
3.84
3.44
41.4
1.9
Hunyuan-3
+0.028
−0.027
3.19
3.25
3.87
3.49
38.7
2.2
GLM-5.2
+0.010
−0.012
3.15
3.20
3.79
3.39
34.5
2.1
Claude-Opus-4.8
−0.014
−0.006
3.09
3.13
3.80
3.39
37.4
4.7
Table 4: UXBench Pro results for seven assistants. Turn BT and Traj. BT are Bradley–Terry strengths derived from pairwise URM judgments. Res. and Aban. denote resolution and abandonment rates. Separability is the fraction of assistant pairs with non-overlapping 95% bootstrap confidence intervals; ρ denotes Spearman correlation.
Formulation
Training signal
S1 ↑
S2 ↑
S3 ↑
S4 ↑
Pooled ↑
Pointwise GRM
Pointwise feedback labels
0.5500
0.5300
0.4650
0.5650
0.5275
General RM
Mixed, general purpose
0.6700
0.5300
0.4850
0.6100
0.5737
BTRM
Regeneration preference
0.6900
0.6100
0.4900
0.5900
0.5950
BTRM
Online dual-answer choice
0.6850
0.6050
0.5400
0.6250
0.6138
Pairwise GRM ✓
Regeneration preference
0.8650
0.8050
0.6200
0.7850
0.7688
Table 5: PairAcc of five trained reward models on the 800-instance URM-Bench under the population condition. Chance is 0.5, exact ties count as errors, and ✓ marks the model selected as the URM for UXBench Pro.
Low effort
High effort
System
Base
+Profile
Base
+Profile
DeepSeek-v4-pro
0.5875
0.5938
0.6175
0.6400
Hunyuan-3
0.6475
0.6625
0.6150
0.6212
GPT-5.5
0.6375
0.6462
0.5863
0.6175
Gemini-3.1-pro
0.6025
0.6225
0.5913
0.5887
Table 7: Thinking-effort ablation on four judges. Values are pooled PairAcc under low and high thinking effort, with and without profile conditioning. The best result for each model is bolded.
Base model
n
S1 ↓
S2 ↓
S3 ↓
S4 ↓
All four ↓
Generic scale
Deepseek-v4-Pro
117
0.248
0.504
0.368
0.376
0.120
Deepseek-v4-Flash
147
0.197
0.327
0.367
0.286
0.082
GPT-5.2
150
0.307
0.573
0.553
0.540
0.260
Hunyuan-3
150
0.833
0.947
0.987
0.953
0.820
Anchored variant
Table 9: State persistence in multi-turn RAG user simulation. Each value is the fraction of sessions in which a state coordinate, or all four jointly, remains unchanged after the first turn. n denotes the number of sessions with at least two scored turns.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Code
Scenario
Assignment signal
WK-a
Office writing
Workplace, business, or professional writing intent
WK-b
Programming
Code intent or programming meta-intent
WK-c
Analysis and decision
Analysis or decision intent; finance for working-role users
WK-d
Professional retrieval
Legal industry
WK-e
Business translation
Translation intent
LF-a
Learning assistance
Mathematics intent, education, or scholastic knowledge
Appendix
Table 10: Deterministic trace-level scenario assignment in UXBench Pro.