Grounding Large Language Models in DSGE Simulators for Policy Generation and Forecasting
Organizations: Birla Institute of Technology and Science, Pilani, India
Abstract
Large language models can produce economic policy responses that sound reasonable, but this does not show that their actions are consistent with economic dynamics. We test this by placing an instruction-tuned language model inside six Snowdrop-backed dynamic stochastic general equilibrium (DSGE) simulators. At each turn, the model observes the economy and a change in economic discourse, selects a bounded policy action, and receives the next simulated state and an economic reward. We implement a common Python interface for repeated rollouts, persistent shocks, state cloning, and rolling-horizon simulation. This setting creates a long-horizon credit-assignment problem. Policy effects may appear several quarters after an action is taken. PPO has a learned value function that can propagate delayed reward to earlier tokens through generalized advantage estimation. GRPO has no learned value function and instead assigns a group-relative advantage from complete rollout returns. It therefore cannot distinguish which earlier turn caused the outcome; if every rollout receives the same return, the normalized advantage is zero. We use PPO as the primary method and GRPO as a matched critic-free baseline. The experiments also test directional semantic signals, reward horizon, trajectory warm starts, cross-simulator transfer, and historically anchored pandemic and monetary-policy shocks. The objective is to judge policy actions by their simulated economic consequences rather than by plausible language alone.
Figures & tables
| Key | Structural model | Principal targets | Policy lever(s) |
|---|---|---|---|
| QPM | Quarterly projection | inflation, output gap | monetary, fiscal |
| SW | Smets–Wouters | inflation, output | monetary, fiscal |
| GSW | Galí–SW pandemic | inflation, output, unemployment | monetary, fiscal |
| IRE | Ireland (2004) | inflation, output gap | monetary |
| MVF | U.S. multivariate filter | inflation, output, unemployment | demand stabilizer |
| RBC | Real business cycle | output, consumption | output stabilizer |
| Source | Contents | Role | Primary metrics |
|---|---|---|---|
| DSGE registry | controlled shocks, states, discourse | training and ablations | return, loss, validity |
| COVID-19 release | scenario shocks and G-Cubed outcomes | structural stress test | RMSE, correlation, loss improvement |
| ECB/EA-MPD | meeting and monthly policy surprises | historical event validation | sign, timing, stabilization loss |
| Model/test | Period | Simulated variables | ||
|---|---|---|---|---|
| QPM, | ||||
| Initial | 16.010000 | 8.461200 | 7.250700 | |
| Next quarter | 13.272388 | 11.004444 | 5.525716 | |
| SW, multiple shocks | ||||
| Initial | 0.000000 | 0.000000 | 0.000000 | |
| Next quarter | 0.020122 | |||
| Horizon | Method | Return | Econ. loss | Valid action | Credit diagnostic |
|---|---|---|---|---|---|
| Short | Base LLM | TBD | TBD | TBD | – |
| Short | GRPO | TBD | TBD | TBD | TBD |
| Short | PPO | TBD | TBD | TBD | TBD |
| Long | Base LLM | TBD | TBD | TBD | – |
| Long | GRPO | TBD | TBD | TBD | TBD |
| Long | PPO | TBD | TBD | TBD | TBD |
| Question | Primary contrast | Hypothesized direction |
|---|---|---|
| Semantic shift | directional vs. static text | higher shock F1; lower loss |
| Reward horizon | rolling vs. one-step | fewer policy-ranking reversals |
| Reward geometry | tolerance band vs. exact target | greater target-zone occupancy |
| Exploration | warm-start vs. direct RL | more useful actions; lower final loss |
| Transfer | multi-model vs. single-model | lower held-out-model loss |
| Scenario / action | Magnitude | Weighting | Base loss | Action loss | Direction | |
|---|---|---|---|---|---|---|
| Recession / cut | 0.05 | Standard | 216 | 219 | worse | |
| Recession / cut | 0.50 | Standard | 216 | 248 | worse | |
| Stagflation / hike | 0.05 | Hawk | 334 | 333 | better | |
| Stagflation / hike | 0.50 | Standard | 369 | 378 | worse | |
| Cost-push / hike | 0.25 | Standard | 2664 | 2996 | worse |