Large language models are increasingly used to simulate how individuals respond to new situations, yet the behavioral reasoning behind these responses is either inherited from pretraining or learned from individual-level annotations, which offer limited behavioral diversity and little supervision of the reasoning itself. We propose to learn behavioral reasoning from prediction markets, whose price trajectories record how populations respond to real-world events at scale. We introduce macro2mind, which trains a language model with GRPO using market signals. A social behavioral decomposition makes behavioral reasoning an explicit step of forecasting: the model infers representative groups of market participants, predicts how each interprets the news and updates its beliefs, reasons about their interactions, and aggregates these responses into a price. A hindsight-regret curriculum with difficulty-aware sampling focuses training on transitions where hindsight-identified groups substantially improve the forecast while prioritizing examples that remain learnable for the current policy. The learned reasoning applies to user simulation without further training. On SWM-Bench, macro2mind achieves state-of-the-art directional accuracy and correlation on Polymarket. Trained on market data, it transfers zero-shot to four user-simulation benchmarks (Humanual, OvertonBench, PRISM, and CAD) and has competitive performance among zero-shot methods. Used as a data generator, macro2mind also raises a downstream simulator's accuracy on unseen users by 15.5 points, outperforming data generated by its backbone by 13.2 points.
Figures & tables
Figure 1: Learning to simulate individuals from macro signals. Left: Each prediction-market price aggregates the behavioral reasoning of its participants: how people interpret an event given their prior beliefs and goals, update those beliefs, and respond, including to one another. Prices record such reasoning at every move, across diverse types of markets. Individual-level data offers only a few answers per person and no record of the reasoning behind them. macro2mind learns behavioral reasoning from market prices and applies it zero-shot to simulate individuals. Right: Relative gains of macro2mind over its Qwen3-4B backbone on market forecasting and zero-shot user simulation. For synthetic data, students are trained on data generated by each model.
Figure 2: The pipeline of macro2mind , illustrated with a running example. (a) Where to learn: The regret-and-difficulty curriculum performs data selection by focusing on high-regret transitions, such as the price jump after inflation data came in lower than expected, while ignoring the noise. (b) How to reason: Through social behavioral decomposition, the model infers how different groups react to the news (first order) and to one another (second order), and aggregates their belief updates into a forecast rewarded by the realized price. (c) Zero-shot transfer: This learned reasoning ability is transferred zero-shot to user simulations, effectively modeling how an individual (e.g., a small-business owner) reacts to the CPI news and peer comments.
Algorithm 1 Overview of macro2mind training with the hindsight-regret curriculum and difficulty-aware sampling
Method
Polymarket
Kalshi
MASE ↓
MAE ↓
DA ↑
Corr ↑
MASE ↓
MAE ↓
DA ↑
Corr ↑
Time-series only
DLinear ( Zeng et al., 2023 )
1.111
0.048
0.580
0.348
1.245
0.075
0.492
− 0.035
TimeMixer ( Wang et al., 2024a )
0.997
0.043
0.590
0.224
1.079
0.065
0.549
− 0.056
PatchTST ( Nie et al., 2022 )
0.994
0.043
0.592
0.306
1.174
0.071
0.536
− 0.035
Time-series + news (Prompting-based LLM)
Table 1: Results on SWM-Bench. Bold / underline : best / second best results.
Humanual
Overton
PRISM
CAD
Method
State ↑
Resp. ↑
MAE ↓
ρ↑
MAE ↓
Acc. ↑
Acc. ↑
Trained user simulators
HumanLM-SFT (Humanual) ( Wu et al., 2026 )
8.2
3.2
0.734
0.566
22.5
59.6
35.0
HumanLM-SFT-think (Humanual)
6.3
1.8
0.754
0.606
23.3
60.1
36.5
HumanLM-GRPO (Humanual)
12.4
5.3
0.718
0.576
21.8
62.6
39.1
Personalized reward models
Table 2: Results on user simulation tasks. Parentheses denote training data ( per bench. : one model trained per benchmark). Bold / underline : best / second best among zero-shot (out-of-domain) results. We multiply all Humanual scores by 100 for reporting.
Figure 6
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Protocol
Users
In-context examples
Out-of-user (main)
100 users
6 known answers of the user
In-user
200 pilot training users
same as in generation requests
Appendix
Table 3: Evaluation protocols for the synthetic data experiment.
Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population's central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience simulation model that targets the individual level within a narrow, well-specified domain. Anacreon learns an authorship embedding that separates individuals, clusters a real qualitative corpus around seed people, and trains a dedicated adapter for each cluster, a mixture of minds, on a Gemma~4 12B base. It harvests demographics, psychological traits, and survey responses from public text, and augments each record with a chain-of-emotion. It reduces prompt brittleness by shuffling response options and reduces positive bias by balancing the training distribution. On a large, externally sourced survey, Anacreon reaches a state-of-the-art ordinal alignment of 0.775, the individual-level accuracy measure on which the field has converged, with a small residual bias. The work is a step toward drawing aggregate insight from faithfully simulated individuals.
Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-derived reason pattern. We study this problem through a 94-person sunscreen concept test in which each respondent evaluated three product concepts and wrote open-ended rationales. We map those rationales into signed reason states Z, where positive signs support adoption and negative signs block it. This gives a practical audit: holding respondent descriptors D, category context K, and concept treatment X fixed, do human rationale-derived reasons help predict behavior Y, and can an LLM simulate the same reason state without seeing the human rationale or outcome? Human rationale-derived reasons substantially improve held-out prediction of purchase intent. LLM-simulated reasons are more brittle: they often sound plausible, but frequently echo the concept board rather than recover the respondent's acceptance or rejection path. The paper contributes an evaluation framework for social simulators. Reason states do not identify natural causal effects by themselves, but they provide an interpretable test of whether a simulator's stated reasons align with human evidence.
Simulating group-level user behavior enables scalable counterfactual evaluation of merchant strategies without costly online experiments. However, building a trustworthy simulator faces two structural challenges. First, information incompleteness causes reasoning-based simulators to over-rationalize when unobserved factors such as offline context and implicit habits are missing. Second, mechanism duality requires capturing both interpretable preferences and implicit statistical regularities, which no single paradigm achieves alone. We propose Policy-Guided Hybrid Simulation (PGHS), a dual-process framework that mines transferable decision policies from behavioral trajectories and uses them as a shared alignment layer. This layer anchors an LLM-based reasoning branch that prevents over-rationalization and an ML-based fitting branch that absorbs implicit regularities. Group-level predictions from both branches are fused for complementary correction. We deploy PGHS on Meituan with 101 merchants and over 26,000 trajectories. PGHS achieves a group simulation error of 8.80%, improving over the best reasoning-based and fitting-based baselines by 45.8% and 40.9% respectively.
Ziyang Chen, Renbing Chen, Daowei Li +4
LongCat Interaction Team, Meituan Inc. Beijing, China · Independent Researcher Changsha, China