cs.CLSep 26, 2025

What Is The Political Content in LLMs' Pre- and Post-Training Data?

Authors: Tanise Ceron, Dmitry Nikolaev, Dominik Stammbach, Debora Nozza

Organizations: Bocconi University, Italy · University of Manchester, UK · Princeton University, USA

Abstract

Large language models (LLMs) reflect politically-slanted opinions in their generated text. Even though it is widely assumed that model behavior stem from training data, there has been no study quantifying the extent to which political content is part of the training data. To bridge this gap, we aim to directly estimate (1)~the proportion of politically engaged texts in training data, (2)~respective data imbalance, (3)~cross-dataset similarity, and (4)~correlations between data composition and model behaviour. We analyze the political content of pre- and post-training datasets of open-source LLMs, combining large-scale sampling, political-leaning classification, and stance detection. We find that all LLM training datasets are systematically skewed towards left-leaning content, with pre-training containing more politically engaged than post-training corpora. We further observe a strong correlation between political stances in training data and model behavior, which is present already in most base models and persists across post-training stages. These findings highlight the role of data composition in correlating with model behavior and motivate the need for greater data transparency as a means to understand and monitor model behavior.

Explore similar work

CardsList
  1. Political Ideology Shifts in Large Language Models

    Aug 22, 2025Pietro Bernardelle, Stefano Civelli, Leon Fröhling +3IdeologyLarge Language Model Bias