From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling
Organizations: The Chinese University of Hong Kong, Shenzhen
Abstract
Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential interaction under limited information. We propose a dynamic bipartite matching framework that combines large language model (LLM) agents with contextual bandits. In a simulated Chinese marriage market, economically grounded LLM agents evaluate locally encountered candidates, while agent-specific Logistic-UCB models learn reciprocal acceptance from realized proposal outcomes. The mechanism therefore separates two decisions---\emph{whom do I like?} and \emph{who is likely to like me back?}---without requiring ex ante market-wide preference rankings. We first validate LLM-induced mate preferences against the empirical conditional-logit reference across multiple LLM backbones. In the matching experiment, Bandit-UCB achieves the highest mean mutual welfare (56.01 versus 54.87 for Gale--Shapley), a smaller gender rank gap than the classical baselines, and the fewest blocking pairs among the LLM-ABM policies. Learned acceptance models show economically interpretable gender-differentiated associations, while counterfactual setups reveal no systematic unilateral advantage from prior search knowledge. Overall, these results support the advantages of decentralized matching with LLM-based behavioral modeling and online learning under incomplete information for economic simulation and computational social science research.
Figures & tables
| Method | Info required | Mutual welfare | Rank gap | #Blocking | Proposals | |
|---|---|---|---|---|---|---|
| Baselines | Gale–Shapley | Whole | 54.87 0.29 | 2.74 0.53 | 0.00 0.00 | 873.14 20.24 |
| LLM-solver | Whole | 54.63 0.50 | 2.57 1.18 | 9.40 10.71 | N/A | |
| Axtell–Kimbrough | Partial | 54.73 0.39 | 2.43 0.83 | 19.90 14.02 | 31504.70 4280.30 | |
| LLM-ABM | Random eligible | Local | 55.24 0.62 | 1.54 0.68 | 38.00 17.42 | 13781.90 2349.29 |
| Utility-only | Local | 55.92 0.48 | 1.84 0.75 | 17.58 12.26 | 16753.14 2876.91 | |
| Bandit-UCB | Local | 56.01 0.59 | 2.11 0.73 | 16.28 12.76 | 16591.78 2376.70 |
| Method | Initial knowledge | Mutual welfare | M rank | F rank | Rank gap | #Blocking | Proposals |
|---|---|---|---|---|---|---|---|
| Gale–Shapley (men-proposing) | Complete rankings | 54.87 0.29 | 16.74 0.39 | 19.48 0.32 | 2.74 0.53 | 0.00 0.00 | 873.14 20.24 |
| Cold-start Bandit-UCB | Local | 55.84 0.33 | 16.48 0.47 | 18.56 0.46 | 2.08 0.70 | 16.70 8.21 | 15829.50 2462.34 |
| All-male warm-start Bandit-UCB | Pooled male | 55.71 0.44 | 16.56 0.73 | 18.68 0.47 | 2.12 0.76 | 15.80 4.66 | 15136.70 2210.78 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Values or scale | Sampling and dependency structure |
|---|---|---|
| Age | 20–50 | Sampled from marriage-market-active age groups according to the target population distribution. |
| Income | Monthly RMB income | Generated conditional on educational attainment around a reference monthly income of 1830 RMB. |
| Education | Middle school or below; high school; university and above | Sampled from population education proportions and aggregated into three experimental categories. |
| Family background | Urban; rural | Sampled from the corresponding urban/rural population composition. |
| Housing | Owns house; no house | Sampled conditionally on age, gender, and family background. |
| Appearance | Below average; average; attractive | Sampled from an approximately bell-shaped distribution, with most probability mass assigned to average appearance. |
| Attribute | Male (n=50) | Female (n=50) | Overall (n=100) |
|---|---|---|---|
| Mean age | 36.60 | 34.62 | 35.61 |
| Mean monthly income | 1881.24 RMB | 1712.88 RMB | 1797.06 RMB |
| Middle school or below | 70% | 72% | 71% |
| High school | 10% | 14% | 12% |
| University and above | 20% | 14% | 17% |
| Urban family background | 58% | 56% | 57% |
| Model | Male | Female | Male WPVR | Female WPVR |
|---|---|---|---|---|
| DeepSeek CoT | 0.692 0.076 | 0.572 0.082 | 0.064 0.025 | 0.122 0.027 |
| Qwen CoT | 0.708 0.115 | 0.646 0.072 | 0.055 0.035 | 0.089 0.022 |
| GPT-OSS CoT | 0.666 0.103 | 0.502 0.092 | 0.073 0.042 | 0.158 0.030 |
| No-CoT baseline | 0.286 0.065 | 0.400 0.085 | 0.298 0.055 | 0.236 0.053 |
| Metric | Male evaluators | Female evaluators |
|---|---|---|
| Kendall’s | 0.703 0.054 | 0.607 0.082 |
| WPVR | 0.054 0.018 | 0.105 0.039 |
| Overlap@3 | 0.460 0.187 | 0.473 0.232 |
| Overlap@5 | 0.684 0.133 | 0.448 0.186 |
| Feature | Interpretation |
|---|---|
| Age gap | Age mismatch between proposer and candidate |
| Income gap | Relative income difference |
| Education gap | Relative education difference |
| Same background | Indicator for shared family background |
| Housing advantage | Relative housing-status advantage |
| Appearance gap | Relative appearance difference |
| Parameter | Value |
|---|---|
| Market size | 50 male agents and 50 female agents |
| Repeated runs | 50 random seeds |
| Main matching LLM scorer | DeepSeek V4 Pro, temperature , CoT prompt |
| Matching horizon | periods |
| Candidate set size | 25 candidates per activation when more are available |
| Reservation utility | on the normalized utility scale |
| Metric | Improved | Unchanged | Worsened | Avg. gain | Avg. loss | Mean effect |
|---|---|---|---|---|---|---|
| Own rank | 11 | 26 | 13 | 0.95 | 0.46 | -0.15 |
| Own score | 11 | 26 | 13 | 0.95 | 0.55 | -0.22 |
| Proposals | 22 | 3 | 25 | 8.20 | 15.37 | -5.00 |