User behavior simulation is the computational modeling of user interactions within information systems through the use of simulated agents in place of live users. It supports system testing and evaluation, decision-making and forecasting, and user experience design. Existing simulators rely on hand-crafted rules or domain expertise that transfers poorly across tasks. SWORD (Simulation-driven Workflow and Prompt Optimization with Role-based Design) is introduced as a framework that jointly optimizes multi-agent workflow topology and natural-language prompts. It is guided solely by a scalar task metric, without domain initialization or task-specific engineering. The experimental results demonstrate that SWORD achieves statistically significant gains over prompt-only, workflow-only, and staged-optimization baselines under a controlled, identical-backbone comparison. Against the strongest published domain-specific baseline, SWORD further improves accuracy while using a smaller backbone model, substantially less training data, and a very reasonable API cost ($4--$6 for each dataset). Beyond predictive performance, SWORD autonomously discovers domain-relevant signals, review-sentiment mapping rules and epidemiological decay priors, purely from scalar error feedback, establishing textual gradients as a mechanism for unsupervised feature-importance discovery in user behavior modeling.
Figures & tables
Figure 1. Comparison of LLM based user simulation paradigms and the SWORD framework. (A) Agent-per-User : instantiates isolated LLM agents per user, preserving personalization but sharing no knowledge. (B) Code-Generation : achieves personalization via generating executable Python simulation code for each user through shared agent workflow, while shifting the burden of workflow design and simulation logic validation to domain experts, limiting task transferability. (C) SWORD (Ours) : is a further step forward compared with Code-Generation to eliminate the dependence on domain experts for different user simulation tasks via automated workflow and prompt optimization and task agnostic modular agent pool design. A three-panel diagram comparing paradigms of user behavior simulation. Panel A shows the Agent-per-User paradigm with isolated LLM agents. Panel B shows the Code-Generation paradigm where an LLM generates a shared Python execution function. Panel C shows the SWORD paradigm featuring a shared workflow DAG (Analyzer, Planner, Generator, Critic, Synthesizer) optimized via an interleaved co-evolution loop of workflow evolution and prompt optimization.
RQ
Addressed in
Evidence
RQ1
§ 6.1 , 6.2
Tables 2 , 3
RQ2
§ 6.5
Table 5
RQ3
§ 6.6
Figure 2 , Table 6
RQ4
§ 6.7
Table 8
Table 1. Research question to section mapping.
Category
Method
Agent MAE ↓
Mask RMSE ↓
Prior work
AI Scientist-v2 Yamada et al.,2025
0.89±0.10
0.45±0.07
YuLan-OneSim Wang et al.,2025a
0.72±0.08
0.42±0.04
G-SIM-ES Holt et al.,2025
0.60±0.07
0.30±0.08
G-SIM-SBI Holt et al.,2025
0.69±0.08
0.20±0.10
Reflexion Shinn et al.,2023
0.57±0.08
0.36±0.04
LLM-based
SOCIA- ∇ (as reported) Hua et al.,2026
0.54±0.06
0.22±0.07
Table 2. Main results on Agent Society (MAE ↓ ) and Mask Adoption (RMSE ↓ ). Results for prior work are sourced from SOCIA- ∇ Hua et al., 2026 . Note: prior-work baselines use GPT-5-class models; SWORD uses GPT-5.4-mini, so this comparison is a cost-efficiency benchmark rather than a controlled accuracy head-to-head. Table 3 provides backbone-controlled ablation comparisons. Best result per column in bold . SWORD results are means over n=10 seeds.
Method
Mask RMSE ↓
Δ [95% CI]
Agent MAE ↓
Δ [95% CI]
Workflow-Only (e.g., Zhang et al., 2025e )
0.0686±0.0722
[−0.004,0.098]
0.456±0.129
[0.061,0.336]∗
Prompt-Only (e.g., Chen et al., 2025c )
0.0232±0.0021
[0.0002,0.0026]∗
0.301±0.087
[−0.040,0.128]
Staged Opt. (e.g., Zhou et al., 2026 )
0.1795±0.1698
[0.036,0.279]∗
0.510±0.274
[0.079,0.427]∗
SWORD (Ours)
0.0217±0.0017
—
0.257±0.105
—
Table 3. Ablation study ( n=10 seeds, Ntrain=30 , Ntest=100 ). All methods share the identical GPT-5.4-mini backbone and identical seeds (paired design). We report mean ± std and the 95% confidence interval of the paired difference vs. SWORD ( Δ , positive means worse than SWORD; undefined for SWORD itself). SWORD’s own 95% CIs of the mean are [0.0206,0.0229] RMSE and [0.182,0.332] MAE (Appendix C ). ∗ marks comparisons significant after per-benchmark Holm–Bonferroni correction ( α=0.05 ). Best per column in bold .
Task
Metric
SWORD
SOCIA- ∇ (reproduced)
agent_society
MAE ↓
0.2567±0.1049
0.2275 ( n=1 )
mask_adoption
RMSE ↓
0.0217±0.0017
0.5547 ( n=1 )
Table 4. SWORD vs. SOCIA- ∇ (our backbone-matched re-execution) on both tasks. SWORD statistics are mean ± std over n=10 independent trials; the re-executed baseline reflects a single successfully executed, non-placeholder run ( n=1 ) due to run-to-run execution instability in its automated code generation.
Seed
Task
Iter
Pre Workflow
Post Workflow
Improvement
1
Agent Society
T0→T1
p→a→s→c→g
p→a→g→t→e
+17.3%
1
Mask Adoption
T1→T2
a→p→g→c→v→s
a→p→g→v→s
+11.9%
2
Agent Society
T2→T3
a→p→g→t→c→v→s
a→p→g→t→v→s
+12.7%
2
Mask Adoption
T0→T1
s→a→g→t→e
a→p→g→c→v→s
+1.2%
3
Agent Society
T2→T3
a→g→s
a→p→g→v→s
+1.9%
3
Mask Adoption
T2→T3
a→p→g→c→v→s
a→p→g→c→t→v→s
+7.8%
Table 5. Rank Inversion Analysis: All 10 seeds ( n=10 ) showing workflow transitions that drive performance improvements. Notation: abbreviated agent names ( a =analyzer, p =planner, g =generator, c =critic, v =validator, s =synthesizer, e =executor, t =transformer) connected by → showing execution order. Improvements are percentage gains in task score (Agent Society: MAE reduction; Mask Adoption: RMSE reduction).
Figure 2. Textual gradient content taxonomy across optimization iterations. Left : Mask Adoption. Role-reframing gradients dominate early and fade as workflow grows more complex, confirming the coarse-to-fine dynamic. Right : Agent Society. 72% of gradients are sentiment-related, autonomously discovered from MAE signals without task-specific initialization. Heat map showing gradient direction category distributions across iterations for both tasks.
Iter
Prompt excerpt
0
“You are a generator agent. Combine all upstream signals into ONE final prediction.”
1
“Priority: (1) historical_rates + recent_trend, (2) intervention status + decay, (3) population priors. Do not omit peer influence or habit dynamics.”
2
“Priority: (1) historical_rates + recent_trend, (2) intervention status + decay, (3) population priors. Do not omit peer influence or habit dynamics.”
3
“If historical_rates empty, apply low-baseline prior (0.08–0.10). In zero-history/no-intervention cases, do not infer high adoption. Intervention effects near-zero before day 10.”
4
“Decision hierarchy: 1) historical/recent trend 2) intervention (exponential decay from start) 3) conservative baseline 4) small population adjustments 5) calibrated params as weak diagnostics only. Clamp to [0, 1].”
Table 6. Generator prompt evolution across iterations on Mask Adoption. Epidemiological domain knowledge (decay, prior, timing) emerges from RMSE signals alone.
Task
T=1
T=2
T=3
T=4
T=5
Mask Adoption
0.8037±0.0095
0.8127±0.0108
0.8217±0.0121
0.8380±0.0145
0.8450±0.0167
Agent Society
0.8950±0.0089
0.8952±0.0091
0.8955±0.0093
0.8958±0.0095
0.8960±0.0097
Table 7. Best-so-far score across n=10 seeds (mean ± std per iteration). Low variance confirms robust, seed-independent convergence by T=5 .
“Apply decay only if applied=true and day ≥ start”
Appendix
Table 9. Gradient direction taxonomy for Mask Adoption. The high proportion of role-reframing gradients (34%) at early iterations confirms that SWORD discovers qualitatively new specialisations, not merely incremental calibrations.
Category
%
Iter
Example direction
Sentiment prioritisation
41
0–1
“Weight review sentiment above numeric priors”
Sentiment mapping
31
1–4
“Strongly positive review: add + 0.8–1.2 stars”
Conflict resolution
16
2–4
“If sentiment disagrees with rating, weight sentiment 2:1”
Calibration
12
3–4
“Platform-mean anchor adjustment: ± 0.1”
Appendix
Table 10. Gradient direction taxonomy for Agent Society. The dominance of sentiment-related gradients (72% combined) shows the textual gradient mechanism correctly identifies review text as the most informative feature.
Seed
SWORD
Staged
Prompt-Only
Workflow-Only
1
0.2445
0.4310
0.3249
0.2965
2
0.3122
0.5078
0.3404
0.3618
3
0.2502
0.7685
0.4649
0.4222
4
0.2000
0.1311
0.2944
0.6161
5
0.1870
0.7424
0.3214
0.4384
6
0.1611
0.2248
0.2562
0.7424
Appendix
Table 11. Per-seed Agent Society MAE ↓ for all methods.
Seed
SWORD
Staged
Prompt-Only
Workflow-Only
1
0.0217
0.2977
0.0218
0.0902
2
0.0201
0.4194
0.0249
0.0254
3
0.0207
0.4194
0.0217
0.0430
4
0.0200
0.0227
0.0230
0.0220
5
0.0202
0.0678
0.0206
0.0394
6
0.0219
0.0200
0.0221
0.0269
Appendix
Table 12. Per-seed Mask Adoption RMSE ↓ for all methods.
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.
Conversational recommender systems (CRSs) are a core component of next-generation intelligent recommender systems because they enable users to actively elicit preferences, clarify intentions, and adapt recommendations in real time. However, there are two key obstacles in the CRS domain: evaluation and access to training data. Evaluating CRSs through real human studies is more critical than for traditional recommender systems, yet such studies are both costly and time-consuming. Moreover, CRS interaction data are often difficult to obtain for model training due to privacy concerns. Large language model (LLM)-based user simulators have shown promise in addressing both challenges by generating synthetic user interactions for evaluation and training. However, existing approaches suffer from systematic positive bias, data leakage, and limited behavioral diversity, and they rely on brittle manual prompt engineering that requires extensive domain expertise. In this paper, we propose a framework to automatically optimize prompts for LLM-based user simulators in CRSs, simultaneously mitigating these issues. Experimental results demonstrate that the proposed framework achieves improved behavioral alignment with human interaction patterns compared to baseline methods across diverse prompt settings.
Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.
Hoang Phan, Dat Huynh, Andrey Zhmoginov +7
Meta Superintelligence Labs · New York University · University of Wisconsin - Madison