Two Americas of Well-Being: Divergent Rural-Urban Patterns of Life Satisfaction and Happiness from 2.6 B Social Media Posts
Authors: Stefano Maria Iacus, Giuseppe Porro
Organizations: Institute for Quantitative Social Science, Harvard University, 1737 Cambridge Street, Cambridge, 100190, Massachusetts, USA. · Department of Law, Economics and Culture, University of Insubria, Via S. Abbondio, 12, Como, 22100, Italy.
Using 2.6 billion geolocated social-media posts (2014-2022) and a fine-tuned generative language model, we construct county-level indicators of life satisfaction and happiness for the United States. We document an apparent rural-urban paradox: rural counties express higher life satisfaction while urban counties exhibit greater happiness. We reconcile this by treating the two as distinct layers of subjective well-being, evaluative vs. hedonic, showing that each maps differently onto place, politics, and time. Republican-leaning areas appear more satisfied in evaluative terms, but partisan gaps in happiness largely flatten outside major metros, indicating context-dependent political effects. Temporal shocks dominate the hedonic layer: happiness falls sharply during 2020-2022, whereas life satisfaction moves more modestly. These patterns are robust across logistic and OLS specifications and align with well-being theory. Interpreted as associations for the population of social-media posts, the results show that large-scale, language-based indicators can resolve conflicting findings about the rural-urban divide by distinguishing the type of well-being expressed, offering a transparent, reproducible complement to traditional surveys.
Figures & tables
Year
N
Correlation (95% CI)
p-value
2014
3143
0.805 [0.792, 0.817]
<10−300
2016
3140
0.496 [0.469, 0.522]
7.5×10−195
2018
3143
0.605 [0.582, 0.627]
3.5×10−313
2020
3143
0.558 [0.533, 0.582]
1.2×10−256
2022
3140
0.442 [0.413, 0.470]
3.5×10−150
Table 1: Correlation between Life Satisfaction and Happiness by Year
Figure 1: Unadjusted county-level means of lifesat and happiness by rurality code (1–9), averaged across 2014–2022. Error bars denote 95% confidence intervals. Life satisfaction increases monotonically with rurality while happiness decreases, documenting the raw rural–urban divergence in expressed well-being prior to any covariate adjustment.
Figure 2: Relationship between Twitter-based life satisfaction and happiness across U.S. counties in 2020. Each point represents a county, colored by its rurality code (1–9) using a diverging RdYlGn color scale, from green (urban, code 1) to red (most remote rural, code 9). The dashed line shows the fitted linear regression with 95% confidence band. The correlation is moderate and positive ( r=0.55 , p<0.001 , N=3,143 ), indicating that, while the two affective indicators are related, they capture distinct facets of expressed well-being.
Figure 3: Relationship between Twitter-based life satisfaction and happiness by rurality code (1–9) in 2020. Each panel shows counties within the same rural classification, with a fitted linear regression (dashed line) and 95% confidence band. Colors follow the same RdYlGn diverging scale as Figure 2 . The positive association between life satisfaction and happiness is consistent across all levels of rurality, though the strength of the correlation varies by group. Notably, panels for codes 8 and 9 display an almost quadratic pattern, reflecting greater heterogeneity in expressed well-being among the most remote rural counties.
Predictor
AME
SE
z
p
lower
upper
acs5 ( × $10,000)
0.0054
0.0019
2.88
0.004
0.0017
0.0091
happiness
1.7623
0.0339
52.01
< 0.001
1.6959
1.8287
margin
-0.0262
0.0100
-2.61
0.009
-0.0459
-0.0065
rural
0.0140
0.0015
9.54
< 0.001
0.0111
0.0168
year = 2016
0.2302
0.0881
2.61
0.009
0.0575
0.4029
year = 2018
0.0257
0.0113
2.27
0.023
0.0035
0.0479
Table 2: Average Marginal Effects for the full lifesat model in equation ( 1 ). Clustered standard errors (HC1) by county via marginaleffects ( Arel-Bundock, 2024 ) .
Figure 4: Marginal effect of partisan vote margin ( margin ) on Pr(lifesat>0) by rurality code, from the full model (Model 4). Points represent average marginal effects; bars denote 95% confidence intervals with clustered standard errors (HC1) by county. The effect is consistently negative across all rurality codes, indicating that Democratic-leaning counties display lower life satisfaction regardless of urban–rural context, though the effect appears strongest in semi-rural counties (codes 6–8) and borderline non-significant at the extremes (codes 1 and 9).
Figure 5: Predicted probability of expressing positive life satisfaction ( P(lifesat>0) ) across the Democratic–Republican vote margin for U.S. counties in 2022, based on the full model in equation ( 1 ). Panels correspond to counties at the 20th, 50th, and 80th percentiles of median household income (ACS 5-year estimates). Lines represent rurality codes (1–9). Life satisfaction decreases as the Democratic vote share increases, with a visually suggestive steeper decline in more-rural counties (higher rural codes); note however that the margin:rural interaction term is not statistically significant under clustered standard errors (see Figure 4 ). Confidence intervals are omitted for visual clarity given the overlap across nine rurality codes; uncertainty estimates are reported in Table 2 .
Figure 6: Predicted probability of expressing positive life satisfaction ( P(lifesat>0) ) across median household income levels for U.S. counties in 2022, based on the full model in equation ( 1 ). Panels correspond to Republican-leaning ( ≤−10 pp), toss-up ( ±10 pp), and Democratic-leaning ( ≥+10 pp) counties. Within each panel the partisan margin is held at the class mean. Lines represent rurality codes (1–9). Life satisfaction increases with household income across all partisan contexts, and the rural–urban gap widens at higher incomes, suggesting that material prosperity amplifies spatial differences in evaluative well-being. Confidence intervals are omitted for visual clarity given the overlap across nine rurality codes; uncertainty estimates are reported in Table 2 .
Predictor
AME
SE
z
p
lower
upper
acs5 ( × $10,000)
-0.0001
0.0007
-0.08
0.936
-0.0014
0.0013
lifesat
0.1409
0.0330
4.27
< 0.001
0.0762
0.2057
margin
-0.0020
0.0096
-0.21
0.831
-0.0208
0.0167
rural
-0.0066
0.0017
-3.95
< 0.001
-0.0099
-0.0033
year = 2016
-0.0703
0.0247
-2.84
0.004
-0.1187
-0.0219
year = 2018
-0.0062
0.0022
-2.77
0.006
-0.0106
-0.0018
Table 3: Average Marginal Effects for the full happiness model. Clustered standard errors (HC1) by county via marginaleffects ( Arel-Bundock, 2024 ) .
Year
State
County
Rural×Margin
Tweets
Poverty
Unemp
Black
Native
BA+
%
%
%
%
%
2014
Alabama
Choctaw
9.000
15168
19.1
5.2
40.0
0.2
13.0
2014
Alabama
Perry
8.000
22413
32.8
15.7
70.8
0.0
15.2
2014
Alabama
Sumter
8.000
38680
30.4
8.2
72.3
0.1
21.8
2014
Alabama
Wilcox
9.000
13189
26.8
10.7
70.2
0.1
11.6
2014
Mississippi
Bolivar
7.000
43579
31.8
7.4
63.2
0.0
27.2
Table 4: Socioeconomic context for extreme counties (ACS 2022 5-year). Rates shown as percentages; last row reports U.S. average.
Predictor
Life Satisfaction
Happiness
Rurality ( rural )
+++
−−−
Partisan Margin ( margin )
− †
0
Rural × Margin
0
0
Household Income ( acs5 )
+ †
0
Year Effects (2016–2022)
+++
−−−
Table 5: Comparative direction and strength of effects in the full models for Life Satisfaction and Happiness.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Dependent variable:
Pr(lifesat>0)
(1)
(2)
(3)
(4)
rural
0.219 ∗∗∗
0.333 ∗∗∗
0.312 ∗∗∗
0.291 ∗∗∗
(0.021)
(0.024)
(0.030)
(0.037)
acs5 ( × $10,000)
0.328 ∗∗∗
0.123 ∗∗
0.120 ∗∗
(0.037)
(0.042)
(0.042)
Appendix
Table 6: Logit estimates of Pr(lifesat>0) by county (2014–2022). All models are weighted by the inverse standard deviation of life satisfaction estimates. Covariates are added sequentially: rurality, partisan margin, income ( acs5 in $10,000), year fixed effects, and their interaction.
Rural.code
AME
SE
z
p
lower
upper
1
-0.0150
0.0081
-1.84
0.066
-0.0309
0.0010
2
-0.0180
0.0074
-2.44
0.015
-0.0324
-0.0036
3
-0.0204
0.0071
-2.85
0.004
-0.0344
-0.0064
4
-0.0268
0.0092
-2.91
0.004
-0.0448
-0.0088
5
-0.0272
0.0102
-2.67
0.008
-0.0472
-0.0072
6
-0.0321
0.0133
-2.41
0.016
-0.0581
-0.0060
Appendix
Table 7: Marginal effect of partisan vote margin ( margin ) on Pr(lifesat>0) by rurality code, from the full model (Model 4). Clustered standard errors (HC1) by county via marginaleffects ( Arel-Bundock, 2024 ) .
Dependent variable:
Pr(happiness>0)
(1)
(2)
(3)
(4)
rural
− 0.420 ∗∗∗
− 0.584 ∗∗∗
− 0.542 ∗∗∗
− 0.500 ∗∗∗
(0.114)
(0.102)
(0.089)
(0.121)
acs5 ( × $10,000)
− 0.402 ∗∗∗
− 0.069
− 0.005
(0.042)
(0.058)
(0.065)
Appendix
Table 8: Weighted logistic regression of above-zero happiness ( P(happiness>0) ) on rurality, partisan vote margin, household income, and year fixed effects.
Figure 7: Predicted probability of expressing positive happiness ( P(happiness>0) ) across the Democratic–Republican vote margin for U.S. counties in 2022, based on the full model in equation ( 2 ). Panels correspond to the 20th, 50th, and 80th percentiles of median household income (ACS 5-year estimates). Lines represent rurality codes (1–9). Happiness levels are uniformly high but decline slightly with rurality, consistently with the (although non statistically significant) margin coefficient in Table 8 . Confidence intervals are omitted for visual clarity given the overlap across nine rurality codes; uncertainty estimates are reported in Table 3 .
Figure 8: Predicted probability of expressing positive happiness ( P(happiness>0) ) across median household income levels for U.S. counties in 2022, based on the full model in equation ( 2 ). Panels correspond to Republican-leaning ( ≤−10 pp), toss-up ( ±10 pp), and Democratic-leaning ( ≥+10 pp) counties. Within each panel the partisan margin is held at the class mean. Lines represent rurality codes (1–9). Happiness increases modestly with income but remains consistently higher in urban counties across all partisan contexts. Political orientation has minimal influence once income and rurality are controlled, consistent with the non-significant margin coefficient in Table 8 and underscoring the predominance of the urban affective advantage over partisan context in shaping hedonic well-being. Confidence intervals are omitted for visual clarity given the overlap across nine rurality codes; uncertainty estimates are reported in Table 3 .
Dependent variable:
lifesat
(1)
(2)
(3)
(4)
rural
0.004 ∗∗∗
0.005 ∗∗∗
0.005 ∗∗∗
0.004 ∗∗∗
(0.0003)
(0.0003)
(0.0003)
(0.0004)
acs5 ( × $10,000)
0.004 ∗∗∗
0.0003
0.0003
(0.0003)
(0.0003)
(0.0003)
Appendix
Table 9: OLS estimates of lifesat by county (2014–2022). All models are weighted by the inverse standard deviation of life satisfaction estimates. Covariates are added sequentially: rurality, partisan margin, income ( acs5 in $10,000), year fixed effects, and their interaction. Clustered standard errors (HC1) by county in parentheses.
Dependent variable:
happiness
(1)
(2)
(3)
(4)
rural
− 0.002 ∗∗∗
− 0.003 ∗∗∗
− 0.002 ∗∗∗
− 0.002 ∗∗∗
(0.0002)
(0.0002)
(0.0002)
(0.0002)
acs5 ( × $10,000)
− 0.002 ∗∗∗
0.001 ∗∗∗
0.001 ∗∗∗
(0.0003)
(0.0002)
(0.0002)
Appendix
Table 10: Linear regression of happiness on rurality, partisan vote margin, household income, and year fixed effects. Results are in line with Table 8 .
Research on political polarization on social media depends on the ability to reliably measure partisanship in user-generated content. However, existing approaches are typically tailored to platform-specific properties, such as structural affordances or linguistic conventions, which hurts generalizability across platforms. This limitation is increasingly consequential as the social media ecosystem fragments and fringe, alt-tech platforms emerge alongside mainstream ones. We propose a text-based, platform-portable methodology for measuring political partisanship in social media posts, anchored by an external news-credibility signal. Posts are embedded using a transformer-based sentence encoder and clustered into topic groups, which are labeled using the aggregated AllSides media bias scores of cited news outlets. A partisanship axis is then constructed in the embedding space as the difference between centroids of oppositely labeled clusters, and individual posts are scored by projection onto this axis. We apply the method to a corpus of approximately 1.3 million posts collected from Bluesky and Truth Social during the six months preceding the 2024 U.S. presidential election, providing the first cross-platform comparison of partisanship distributions on these two ideologically asymmetric platforms. The resulting partisanship scores correlate significantly with held-out AllSides media bias scores both in-distribution and out-of-distribution on an independent Twitter corpus, and recover within-platform partisan dynamics that platform identity alone cannot explain.
Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs). We introduce UrbanWell, a large-scale benchmark designed to systematically evaluate the spatio-temporal reasoning capabilities of MLLMs for urban wellbeing analytics through joint modeling of satellite and street view imagery. UrbanWell spans 38 cities across multiple years and includes diverse indicators covering (1) environmental conditions (CO2, NO2, PM2.5, and Normalized Difference Vegetation Index), (2) spatial accessibility (minimum distance to supermarkets and restaurants), (3) urban form (road length, road density, and land use), (4) urban vitality (population, economic activity diversity, and land use diversity), and (5) subjective perception attributes (e.g., safety, beauty, liveliness, wealth, and quietness). All indicators are aligned at grid level to enable standardized evaluation. Beyond static prediction, UrbanWell defines temporal reasoning tasks, including future value forecasting from historical observations and temporal trend classification. We benchmark 15 state-of-the-art representative MLLMs in a zero-shot setting, providing a comprehensive comparative evaluation across spatial and temporal dimensions. Experimental results indicate that while MLLMs capture salient spatial and perceptual cues, their performance varies substantially across heterogeneous urban indicators spanning environment and subjective perception. UrbanWell serves as a unified benchmark for evaluating multimodal spatial and temporal reasoning in urban wellbeing analytics, offering a standardized testbed for systematic assessment and future research on multimodal urban intelligence. Our codes and datasets are accessible via https://github.com/axin1301/UrbanWell-Benchmark.
Yanxin Xi, Xiang Su, Jie Feng +3
Department of Computer Science, University of Helsinki, Helsinki, Finland · Department of Computer Science, Department of Agricultural Sciences, University of Helsinki, Helsinki, Finland · Zhongguancun Academy, Beijing, China +2
Political polarization has become a defining feature of online discourse, yet its long-term evolution remains poorly understood. We present a longitudinal analysis of ideological polarization in Reddit discussions by measuring semantic differences in the language used by opposing political communities. We construct temporally aligned community-specific word embeddings and quantify ideological polarization as the semantic divergence of political concepts over time. Our analysis shows that ideological polarization has increased substantially during the study period, both at the concept- and topic-level. Unlike prior computational work, which has largely focused on cross-sectional analyses or affective dimensions of polarization at a single point at time, our approach captures the evolution of ideological differences in semantic framing. The proposed framework provides a scalable method for studying the temporal dynamics of ideological polarization in large-scale social media discourse.