This paper demonstrates how reinforcement learning can explain two puzzling empirical patterns in household consumption behavior during economic downturns. I develop a model where agents use Q-learning with neural network approximation to make consumption-savings decisions under income uncertainty, departing from standard rational expectations assumptions. The model replicates two key findings from recent literature: (1) unemployed households with previously low liquid assets exhibit substantially higher marginal propensities to consume (MPCs) out of stimulus transfers compared to high-asset households (0.50 vs 0.34), even when neither group faces borrowing constraints, consistent with Ganong et al. (2024); and (2) households with more past unemployment experiences maintain persistently lower consumption levels after controlling for current economic conditions, a "scarring" effect documented by Malmendier and Shen (2024). Unlike existing explanations based on belief updating about income risk or ex-ante heterogeneity, the reinforcement learning mechanism generates both higher MPCs and lower consumption levels simultaneously through value function approximation errors that evolve with experience. Simulation results closely match the empirical estimates, suggesting that adaptive learning through reinforcement learning provides a unifying framework for understanding how past experiences shape current consumption behavior beyond what current economic conditions would predict.
Figures & tables
Parameter
Value
Notes
Simulation Parameters
number of agents
50
number of time periods
50
learning rate (simulation)
1.1×10−3
optimizer (simulation)
ADAM
less sensitive to learning-rate choice
polynomial degree
5
Calibrated, time 0 fit
Table 1: Key Configuration Parameters
Figure 1: Consumption policies at t=5 for the experiment with repeated employment with one agent receiving a series of unemployment realizations and the other receiving a series of employment realizations. The blue line is the agent with unemployment experiences, while the red line is the agent with employment experiences. The dotted gray line is the rational agent’s policy. The figure on the left represents the agent’s consumption policy as a function of their assets (x-axis) and income while employed ye , while on the right it represents the agent’s consumption policy as a function of their assets and income while unemployed yu . Value functions use a polynomial fit.
Figure 2: MPCs for the experiment with repeated employment and unemployment realizations after t=5 periods, where one agent receives a series of unemployment realizations and the other receives a series of employment realizations. The blue line is the agent with unemployment experiences, while the red line is the agent with employment experiences. The dotted gray line is the rational benchmark’s policy. The figure on the left represents the agent’s marginal propensity to consume as a function of their assets (x-axis) and income while employed ye , while on the left it represents the agent’s consumption policy as a function of their assets and income while unemployed yu . This uses the polynomial fit to the value function generating local oscillations from the polynomial.
Figure 3: Neural network gradient updates produce larger changes to EV at savings levels closer to the agent’s current position. This increases the slope and curvature of EV , which in turn decreases average consumption and increases MPCs out of a one-time transfer.
Figure 4: EV evaluated at the agent’s current asset levels before (blue) and after (red) a single gradient update at t=0 following 5 periods of unemployment. The x-axis is next-period assets a′ . The red curve steepens (lower consumption, higher savings) and becomes more concave (higher MPCs).
Figure 5: EV before (blue) and after (red) the 4th gradient update. The pattern of steepening and increased curvature persists, consistent with the cumulative mechanism.
Figure 6: Consumption policy as a function of assets for different values of slope ( g1 ) and curvature ( g2 ) parameters in the polynomial EV specification Eq. 15 . Higher slope reduces consumption everywhere (scarring). Higher curvature raises consumption at high asset values.
Figure 7: MPCs as a function of assets for different values of slope ( g1 ) and curvature ( g2 ) in the polynomial EV . Higher curvature substantially increases MPCs everywhere. Higher slope slightly decreases MPCs.
Liquidity Type
Model MPC
Empirical MPC a
Rational MPC Estimates b
Low
0.501
0.53
0.03 (1-asset), 0.16 (2-asset)
High
0.343
0.29
0.03 (1-asset), 0.16 (2-asset)
Table 2: Model versus Empirical Marginal Propensities to Consume (MPCs)
Dependent variable:
consumption
(1)
(2)
UEP t
− 0.0884 ∗∗∗
-0.0378 ∗∗∗
( − 0.108, − 0.069)
( − 0.052, − 0.024)
assets t
0.1265 ∗∗∗
(0.123, 0.130)
Table 4: Regression of personal unemployment experience on consumption with and without controls for asset levels.
Model
Consumption
MPC
Empirical
↓
↑
RL Agent
↓✓
↑✓
Pessimistic (AU)
↓✓
↓×
Rational (benchmark)
−
−
Table 5: Directional Comparison: RL vs. Pessimistic Beliefs vs. Rational Expectations
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Consumption as a function of assets and income after 50 periods of learning. No polynomial smoothing is used. The blue line is the reinforcement learner, the orange line is the rational benchmark. 95th percentiles are plotted.
Figure A2: MPC as a function of assets and income after 50 periods of learning. No polynomial smoothing is used. The blue line shows the reinforcement learner, the orange line shows the rational benchmark. 95th percentile intervals are plotted.
Figure A3: Consumption policies at t=0 for the experiment with repeated employment with one agent receiving a series of unemployment realizations and the other receiving a series of employment realizations. The blue line is the agent with unemployment experiences, while the red line is the agent with employment experiences. The dotted gray line is the rational agent’s policy. The figure on the left represents the agent’s consumption policy as a function of their assets (x-axis) and income while employed ye , while on the right it represents the agent’s consumption policy as a function of their assets and income while unemployed yu .
Figure A4: Orange line is rational agent solution under 1.5x pessimistic income transition probabilities, blue line is rational agent solution under original income transition probabilities. The figure on the left represents the agent’s consumption policy as a function of their assets (x-axis) and income while unemployed yu , while on the right it represents the agent’s consumption policy as a function of their assets and income while unemployed ye .
Figure A5: Orange line is rational agent marginal propensity to consume under 1.5x pessimistic income transition probabilities, blue line is rational agent solution under original income transition probabilities. The figure on the left represents the agent’s consumption policy as a function of their assets (x-axis) and income while unemployed yu , while on the right it represents the agent’s consumption policy as a function of their assets and income while unemployed ye .
Figure A6: Initial consumption policy fit at t=0 quarters for a single seed, no smoothing.
Figure A7: Consumption policy and MPC fit against rational at t=10 quarters for a single seed, no smoothing, systematic tilting of policy function away from rational, as well as local fluctuations without smoothing.
Figure A8: Consumption policy and MPC fit against rational at t=240 quarters for a single seed, no smoothing. Local fluctuations remain elevated but systematic bias in policy has mostly vanished.