cs.AISep 29, 2026

Character Training for Risk-Averse Agents

Authors: Arav Dhoot, Punya Syon Pandey, Jamie Johnson, Daniel Tan, Elliott Thornley, David Demitri Africa

Organizations: UK AI Security Institute · Resolution

Abstract

Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents.

Figures & tables

Appendix figures & tables28 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Out-of-Distribution Generalization of Risk Aversion in Language Models

    Jul 2, 2026Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie +3RiskLarge Language Model Safety

  2. AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

    Jun 4, 2025Akshat Naik, Emma Gouné, Patrick Quinn +4Large Language Model AgentsMisalignment Persona

  3. Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors

    May 7, 2026Jonas Wiedermann-Möller, Leonard Dung, Maksym AndriushchenkoArtificial Intelligence SafetyLarge Language Model Agents