cs.CLFeb 11, 2026

Evaluating Alignment of Behavioral Dispositions in LLMs

Authors: Amir Taubenfeld, Zorik Gekhman, Lior Nezry, Omri Feldman, Natalie Harris, Shashir Reddy, Romina Stella, Ariel Goldstein, +3 more

Organizations: Google Research · Hebrew University · University of Cambridge

Abstract

As people turn to LLMs for social advice, understanding their behavior in such contexts becomes essential. In this work, we focus on behavioral dispositions: the underlying tendencies that shape responses in social contexts. We introduce STAR, a framework for studying how closely the dispositions expressed by LLMs align with those of humans. STAR builds on established psychological questionnaires, adapting their items into realistic advice-seeking scenarios, as self-report may not transfer to actual advisory behavior. Using STAR, we construct a dataset of 23k scenarios, each validated by 3 raters and annotated with preferences from 10 participants. Across 25 LLMs, we find that (1) when human consensus is high, frontier models can fail to reflect it in 15-20% of cases, and smaller models fail at substantially higher rates; (2) when humans disagree, LLM recommendations are substantially less diverse than human choices, both within individual models and even across models from different providers, potentially narrowing the range of options users are guided toward; (3) LLMs' self-reported values are poor predictors of their recommendations. To support future research we make our dataset and code publicly available.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

    Jun 10, 2026Rafal Kocielnik, Pengrui Han, Peiyang Song +5PersonalityPsychometric Properties

  2. Teaching Values to Machines: Simulating Human-Like Behavior in LLMs

    May 28, 2026Asaf Yehudai, Naama Rozen, Ariel GeraHuman ValuesUser Simulation

  3. LLMs Can Better Capture Human Judgments--With the Right Prompts

    Jun 10, 2026Danica Dillion, Chen Cecilia Liu, Baihui Wang +5Human JudgmentMoral Reasoning