Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices
Organizations: DIMES Dept., University of Calabria, Italy
Abstract
Evaluations of Large Language Models (LLMs) morality typically consider decisions in isolation, thus overlooking whether an individual's unrelated prior conduct influences the model's subsequent choices. This leaves open the question of whether, and to what extent, moral history shapes LLM decisional behaviors. Prior work on human moral decision-making shows that past behavior can influence subsequent moral choices. Building on this observation, we investigate whether analogous effects emerge in LLMs in two complementary ways: at the behavioral level, through the model's observable responses, and at the representation level, through its latent internal representations. We introduce MoralLedger, a framework for studying how an actor's moral history shapes actions for LLMs' behaviors under a fixed decision context. At the behavioral level, we find that prior moral histories systematically alter subsequent choices as a function of their valence and intensity. At the internal representation level, these histories induce a linearly recoverable direction in the residual stream that generalizes to held-out examples. Intervening along this direction on neutral-history prompts produces two-sided intensity-dependent changes in subsequent choices, with effects that are stronger than those induced by prompting alone or by favorable-nonmoral direction. To our knowledge, this is the first demonstration that a latent representation of an actor's prior moral conduct can provide signed inference-time control over a moral decision. Our MoralLedger extends moral evaluation beyond static dilemmas, establishing moral history as both a source of behavioral sensitivity and a causal target for auditing and controlling moral behavior in LLMs.
Figures & tables
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbols | Meaning |
|---|---|
| Moral domains, one domain, and its three-valence history block. | |
| History with valence , decision context, and its three available actions. | |
| Valence set; indexes prior history, while indexes a current action. | |
| Disjoint history–context pairs for direction estimation, validation selection, and evaluation. | |
| Single-token answer labels, one label, rotation index, and rotation operator. | |
| Tokenized prompt under history and rotation ; next-token logit for label . |
| Component | Source | Validation | Role |
|---|---|---|---|
| Moral histories | Social Chemistry 101 ( Forbes et al., 2020 ) | Human action-valence annotations | Assigned history treatment |
| Decision contexts | Controlled construction | Author review, arithmetic checks, LLM-as-a-judge | Three-action choice |
| Favorable histories | crowd-enVENT ( Troiano et al., 2023 ) | Human appraisals and controlled neutral counterparts | Valence control |
| External actions | Moral Stories ( Emelin et al., 2021 ) | Human-authored stories and labeled actions | Transfer evaluation |
| Domain | Decision families | Moral endpoint |
|---|---|---|
| Care/helping | Prevent loss; material aid; time aid | Maximize help |
| Fairness/entitlement | Credit; opportunities; resources | Minimize own share |
| Honesty/property | Transaction error; claim; disclosure | Return/disclose accurately |
| Responsibility/effort | Obligation; duty; workload | Maximize accepted share |
| Model | History contrast | [95% CI] | (%) [95% CI] |
|---|---|---|---|
| Gemma 2 2B | Positive neutral | +0.296 [-0.020, +0.676] | +3.2 [+0.1, +6.6] ∗ |
| Negative neutral | -0.273 [-0.635, +0.094] | -2.2 [-6.2, +2.0] | |
| Positive negative | +0.569 [+0.194, +0.979] ∗ | +5.3 [+0.9, +10.2] ∗ | |
| Gemma 2 9B | Positive neutral | +1.044 [+0.389, +1.767] ∗ | +9.7 [+2.6, +17.8] ∗ |
| Negative neutral | -0.610 [-1.449, +0.109] | -5.6 [-13.8, +1.5] | |
| Positive negative | +1.654 [+0.858, +2.539] ∗ | +15.4 [+6.8, +24.6] ∗ |
| Contrast | Model | Estimate [95% CI] | |
|---|---|---|---|
| Construct: is the effect about moral history? | |||
| moral (pos) non-moral (pos) (2/8) | Gemma 2 2B | +0.217 [-0.110, +0.592] | .224 |
| Gemma 2 9B | +0.638 [-0.042, +1.440] | .088 | |
| Llama 3.2 3B | +0.128 [-0.028, +0.329] | .150 | |
| Llama 3.1 8B | +0.167 [-0.098, +0.464] | .220 | |
| OLMo 2 1B | +0.040 [+0.007, +0.071] | .019 ∗ | |
| Coordinate | At that coordinate | Best layer | |||||
| Model | Block | Depth | Position | AUC [95% CI] | Paired rank | AUC | Diff |
| Gemma 2 2B | 11 | 42% | history | 0.830 [0.696, 0.937] | 0.825 | 0.914 | -0.085 |
| Gemma 2 9B | 19 | 45% | history | 0.936 [0.856, 0.986] | 0.925 | 0.936 | +0.000 |
| Llama 3.2 3B | 27 | 96% | decision | 0.684 [0.609, 0.768] | 0.800 | 0.756 | -0.073 |
| Llama 3.1 8B | 21 | 66% | decision | 0.739 [0.643, 0.838] | 0.783 | 0.741 | -0.002 |
| OLMo 2 1B | 3 | 19% | decision | 0.585 [0.506, 0.666] | 0.708 | 0.608 | -0.023 |
| Probability change [95% CI] (%) | ||||||
|---|---|---|---|---|---|---|
| Model | at | at | Both | Validity | Prompt | |
| Gemma 2 2B | 0.15 | +22.4 [+14.4, +30.5] | -11.1 [-18.0, -5.1] | ✓ | ✓ | ✓ |
| Gemma 2 9B | 0.25 | +20.6 [+10.2, +31.6] | -25.6 [-34.9, -16.9] | ✓ | ✓ | ✓ |
| Llama 3.2 3B | 0.20 | +5.7 [+4.1, +7.4] | -5.3 [-6.9, -3.7] | ✓ | – | – |
| Llama 3.1 8B | 0.25 | +7.0 [+3.2, +11.2] | -12.1 [-18.0, -6.8] | ✓ | ✓ | – |
| OLMo 2 1B | 0.10 | +1.9 [+1.2, +2.7] | -1.5 [-2.1, -1.0] | ✓ | ✓ | – |
| Model | |||
|---|---|---|---|
| Gemma 2 2B | +0.224 | +1.942 | 8.7 |
| Gemma 2 9B | +0.709 | +1.739 | 2.5 |
| Llama 3.2 3B | +0.105 | +0.154 | 1.5 |
| Llama 3.1 8B | +0.088 | +0.720 | – |
| OLMo 2 1B | +0.007 | +0.039 | – |
| OLMo 2 7B | +0.109 | +0.303 | 2.8 |
| Log-odds specificity | Probability specificity (%) | |||
|---|---|---|---|---|
| Model | [95% CI] | [95% CI] | [95% CI] | [95% CI] |
| Gemma 2 2B | +1.702 [+1.230, +2.218] ∗ | +1.224 [+0.581, +1.915] ∗ | +23.4 [+17.1, +30.3] ∗ | +9.6 [+2.2, +18.2] ∗ |
| Gemma 2 9B | +1.892 [+1.390, +2.464] ∗ | +1.667 [+1.093, +2.287] ∗ | +17.3 [+10.2, +25.5] ∗ | +15.7 [+9.2, +23.1] ∗ |
| Llama 3.2 3B | +0.021 [+0.008, +0.033] ∗ | +0.053 [+0.035, +0.075] ∗ | +0.5 [+0.2, +0.8] ∗ | +0.9 [+0.5, +1.2] ∗ |
| Llama 3.1 8B | +0.118 [-0.014, +0.279] | +0.542 [+0.102, +1.087] ∗ | +3.0 [+1.2, +5.1] ∗ | +5.7 [+3.1, +8.4] ∗ |
| OLMo 2 1B | +0.302 [+0.230, +0.395] ∗ | +0.175 [+0.122, +0.244] ∗ | +4.2 [+2.9, +5.8] ∗ | +3.3 [+2.4, +4.4] ∗ |
| Model | Negative steering | Positive steering |
|---|---|---|
| Gemma 2 2B | ||
| Gemma 2 9B | ||
| Llama 3.2 3B | ||
| Llama 3.1 8B | ||
| OLMo 2 1B | ||
| OLMo 2 7B |
| Log odds | |||
|---|---|---|---|
| Model | Sign | Transfer [95% CI] | Moral minus favorable [95% CI] |
| Gemma 2 2B | +0.698 [+0.429, +0.964] ∗ | +0.701 [+0.535, +0.868] ∗ | |
| -0.146 [-0.371, +0.063] | -0.520 [-0.711, -0.344] ∗ | ||
| Gemma 2 9B | +1.149 [+0.954, +1.354] ∗ | +1.091 [+0.938, +1.243] ∗ | |
| -1.806 [-2.030, -1.587] ∗ | -1.281 [-1.479, -1.098] ∗ | ||
| Llama 3.2 3B | +0.455 [+0.395, +0.512] ∗ | +0.097 [+0.077, +0.119] ∗ | |
| Negative steering, | ||||||
|---|---|---|---|---|---|---|
| Model | MMLU | ARC | WinoGrande | GSM8K | HellaSwag | TruthfulQA |
| Gemma 2 2B | -2.3 | -1.5 | -0.7 | -12.8 | -1.4 | -2.6 |
| Gemma 2 9B | -2.6 | -1.2 | -1.3 | +0.7 | +0.1 | +7.0 |
| Llama 3.2 3B | -1.3 | +1.1 | -5.4 | +1.1 | +0.5 | +0.0 |
| Llama 3.1 8B | +1.0 | -4.5 | -3.9 | +0.5 | -0.9 | +2.5 |
| OLMo 2 1B | +4.0 | +0.0 | -1.6 | -6.5 | +5.7 | +5.7 |
| Same vector as main experiments | Random | Bench.-calibrated | |||
|---|---|---|---|---|---|
| Model | at | ||||
| Gemma 2 2B | -2.66 | -3.17 | 0.15 | -1.02 | -0.68 |
| Gemma 2 9B | -1.02 | -7.75 | 0.25 | -0.93 | -0.27 |
| Llama 3.2 3B | -0.65 | -1.79 | 0.20 | -1.19 | -1.17 |
| Llama 3.1 8B | -0.72 | -2.30 | 0.25 | -0.86 | +0.55 |
| OLMo 2 1B | +1.48 | +1.48 | 0.10 | +1.21 | +1.71 |