User behavior simulation is the computational modeling of user interactions within information systems through the use of simulated agents in place of live users. It supports system testing and evaluation, decision-making and forecasting, and user experience design. Existing simulators rely on hand-crafted rules or domain expertise that transfers poorly across tasks. SWORD (Simulation-driven Workflow and Prompt Optimization with Role-based Design) is introduced as a framework that jointly optimizes multi-agent workflow topology and natural-language prompts. It is guided solely by a scalar task metric, without domain initialization or task-specific engineering. The experimental results demonstrate that SWORD achieves statistically significant gains over prompt-only, workflow-only, and staged-optimization baselines under a controlled, identical-backbone comparison. Against the strongest published domain-specific baseline, SWORD further improves accuracy while using a smaller backbone model, substantially less training data, and a very reasonable API cost ($4--$6 for each dataset). Beyond predictive performance, SWORD autonomously discovers domain-relevant signals, review-sentiment mapping rules and epidemiological decay priors, purely from scalar error feedback, establishing textual gradients as a mechanism for unsupervised feature-importance discovery in user behavior modeling.
Figures & tables
Figure 1. Comparison of LLM based user simulation paradigms and the SWORD framework. (A) Agent-per-User : instantiates isolated LLM agents per user, preserving personalization but sharing no knowledge. (B) Code-Generation : achieves personalization via generating executable Python simulation code for each user through shared agent workflow, while shifting the burden of workflow design and simulation logic validation to domain experts, limiting task transferability. (C) SWORD (Ours) : is a further step forward compared with Code-Generation to eliminate the dependence on domain experts for different user simulation tasks via automated workflow and prompt optimization and task agnostic modular agent pool design. A three-panel diagram comparing paradigms of user behavior simulation. Panel A shows the Agent-per-User paradigm with isolated LLM agents. Panel B shows the Code-Generation paradigm where an LLM generates a shared Python execution function. Panel C shows the SWORD paradigm featuring a shared workflow DAG (Analyzer, Planner, Generator, Critic, Synthesizer) optimized via an interleaved co-evolution loop of workflow evolution and prompt optimization.
RQ
Addressed in
Evidence
RQ1
§ 6.1 , 6.2
Tables 2 , 3
RQ2
§ 6.5
Table 5
RQ3
§ 6.6
Figure 2 , Table 6
RQ4
§ 6.7
Table 8
Table 1. Research question to section mapping.
Category
Method
Agent MAE ↓
Mask RMSE ↓
Prior work
AI Scientist-v2 Yamada et al.,2025
0.89±0.10
0.45±0.07
YuLan-OneSim Wang et al.,2025a
0.72±0.08
0.42±0.04
G-SIM-ES Holt et al.,2025
0.60±0.07
0.30±0.08
G-SIM-SBI Holt et al.,2025
0.69±0.08
0.20±0.10
Reflexion Shinn et al.,2023
0.57±0.08
0.36±0.04
LLM-based
SOCIA- ∇ (as reported) Hua et al.,2026
0.54±0.06
0.22±0.07
Table 2. Main results on Agent Society (MAE ↓ ) and Mask Adoption (RMSE ↓ ). Results for prior work are sourced from SOCIA- ∇ Hua et al., 2026 . Note: prior-work baselines use GPT-5-class models; SWORD uses GPT-5.4-mini, so this comparison is a cost-efficiency benchmark rather than a controlled accuracy head-to-head. Table 3 provides backbone-controlled ablation comparisons. Best result per column in bold . SWORD results are means over n=10 seeds.
Method
Mask RMSE ↓
Δ [95% CI]
Agent MAE ↓
Δ [95% CI]
Workflow-Only (e.g., Zhang et al., 2025e )
0.0686±0.0722
[−0.004,0.098]
0.456±0.129
[0.061,0.336]∗
Prompt-Only (e.g., Chen et al., 2025c )
0.0232±0.0021
[0.0002,0.0026]∗
0.301±0.087
[−0.040,0.128]
Staged Opt. (e.g., Zhou et al., 2026 )
0.1795±0.1698
[0.036,0.279]∗
0.510±0.274
[0.079,0.427]∗
SWORD (Ours)
0.0217±0.0017
—
0.257±0.105
—
Table 3. Ablation study ( n=10 seeds, Ntrain=30 , Ntest=100 ). All methods share the identical GPT-5.4-mini backbone and identical seeds (paired design). We report mean ± std and the 95% confidence interval of the paired difference vs. SWORD ( Δ , positive means worse than SWORD; undefined for SWORD itself). SWORD’s own 95% CIs of the mean are [0.0206,0.0229] RMSE and [0.182,0.332] MAE (Appendix C ). ∗ marks comparisons significant after per-benchmark Holm–Bonferroni correction ( α=0.05 ). Best per column in bold .
Task
Metric
SWORD
SOCIA- ∇ (reproduced)
agent_society
MAE ↓
0.2567±0.1049
0.2275 ( n=1 )
mask_adoption
RMSE ↓
0.0217±0.0017
0.5547 ( n=1 )
Table 4. SWORD vs. SOCIA- ∇ (our backbone-matched re-execution) on both tasks. SWORD statistics are mean ± std over n=10 independent trials; the re-executed baseline reflects a single successfully executed, non-placeholder run ( n=1 ) due to run-to-run execution instability in its automated code generation.
Seed
Task
Iter
Pre Workflow
Post Workflow
Improvement
1
Agent Society
T0→T1
p→a→s→c→g
p→a→g→t→e
+17.3%
1
Mask Adoption
T1→T2
a→p→g→c→v→s
a→p→g→v→s
+11.9%
2
Agent Society
T2→T3
a→p→g→t→c→v→s
a→p→g→t→v→s
+12.7%
2
Mask Adoption
T0→T1
s→a→g→t→e
a→p→g→c→v→s
+1.2%
3
Agent Society
T2→T3
a→g→s
a→p→g→v→s
+1.9%
3
Mask Adoption
T2→T3
a→p→g→c→v→s
a→p→g→c→t→v→s
+7.8%
Table 5. Rank Inversion Analysis: All 10 seeds ( n=10 ) showing workflow transitions that drive performance improvements. Notation: abbreviated agent names ( a =analyzer, p =planner, g =generator, c =critic, v =validator, s =synthesizer, e =executor, t =transformer) connected by → showing execution order. Improvements are percentage gains in task score (Agent Society: MAE reduction; Mask Adoption: RMSE reduction).
Figure 2. Textual gradient content taxonomy across optimization iterations. Left : Mask Adoption. Role-reframing gradients dominate early and fade as workflow grows more complex, confirming the coarse-to-fine dynamic. Right : Agent Society. 72% of gradients are sentiment-related, autonomously discovered from MAE signals without task-specific initialization. Heat map showing gradient direction category distributions across iterations for both tasks.
Iter
Prompt excerpt
0
“You are a generator agent. Combine all upstream signals into ONE final prediction.”
1
“Priority: (1) historical_rates + recent_trend, (2) intervention status + decay, (3) population priors. Do not omit peer influence or habit dynamics.”
2
“Priority: (1) historical_rates + recent_trend, (2) intervention status + decay, (3) population priors. Do not omit peer influence or habit dynamics.”
3
“If historical_rates empty, apply low-baseline prior (0.08–0.10). In zero-history/no-intervention cases, do not infer high adoption. Intervention effects near-zero before day 10.”
4
“Decision hierarchy: 1) historical/recent trend 2) intervention (exponential decay from start) 3) conservative baseline 4) small population adjustments 5) calibrated params as weak diagnostics only. Clamp to [0, 1].”
Table 6. Generator prompt evolution across iterations on Mask Adoption. Epidemiological domain knowledge (decay, prior, timing) emerges from RMSE signals alone.
Task
T=1
T=2
T=3
T=4
T=5
Mask Adoption
0.8037±0.0095
0.8127±0.0108
0.8217±0.0121
0.8380±0.0145
0.8450±0.0167
Agent Society
0.8950±0.0089
0.8952±0.0091
0.8955±0.0093
0.8958±0.0095
0.8960±0.0097
Table 7. Best-so-far score across n=10 seeds (mean ± std per iteration). Low variance confirms robust, seed-independent convergence by T=5 .
“Apply decay only if applied=true and day ≥ start”
Appendix
Table 9. Gradient direction taxonomy for Mask Adoption. The high proportion of role-reframing gradients (34%) at early iterations confirms that SWORD discovers qualitatively new specialisations, not merely incremental calibrations.
Category
%
Iter
Example direction
Sentiment prioritisation
41
0–1
“Weight review sentiment above numeric priors”
Sentiment mapping
31
1–4
“Strongly positive review: add + 0.8–1.2 stars”
Conflict resolution
16
2–4
“If sentiment disagrees with rating, weight sentiment 2:1”
Calibration
12
3–4
“Platform-mean anchor adjustment: ± 0.1”
Appendix
Table 10. Gradient direction taxonomy for Agent Society. The dominance of sentiment-related gradients (72% combined) shows the textual gradient mechanism correctly identifies review text as the most informative feature.
Seed
SWORD
Staged
Prompt-Only
Workflow-Only
1
0.2445
0.4310
0.3249
0.2965
2
0.3122
0.5078
0.3404
0.3618
3
0.2502
0.7685
0.4649
0.4222
4
0.2000
0.1311
0.2944
0.6161
5
0.1870
0.7424
0.3214
0.4384
6
0.1611
0.2248
0.2562
0.7424
Appendix
Table 11. Per-seed Agent Society MAE ↓ for all methods.
Seed
SWORD
Staged
Prompt-Only
Workflow-Only
1
0.0217
0.2977
0.0218
0.0902
2
0.0201
0.4194
0.0249
0.0254
3
0.0207
0.4194
0.0217
0.0430
4
0.0200
0.0227
0.0230
0.0220
5
0.0202
0.0678
0.0206
0.0394
6
0.0219
0.0200
0.0221
0.0269
Appendix
Table 12. Per-seed Mask Adoption RMSE ↓ for all methods.