Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
Organizations: Indiana University Bloomington / United States
Abstract
Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across reasoning steps remains underexplored. We introduce moral reasoning trajectories, sequences of ethical framework invocations across intermediate reasoning steps, and analyze their dynamics across six models and three benchmarks. We find that moral reasoning involves systematic multi-framework deliberation: 55.4--57.7% of consecutive steps involve framework switches, and only 16.4--17.8% of trajectories remain framework-consistent. Unstable trajectories remain 1.29 times more susceptible to persuasive attacks (p=0.015). At the representation level, linear probes localize framework-specific encoding to model-specific layers (layer 63/81 for Llama-3.3-70B; layer 17/81 for Qwen2.5-72B), achieving 16.8--22.2% lower KL divergence than the step-prior baseline. Activation steering applied during generation moves the framework-consistency--accuracy relationship, widening it for Qwen2.5-72B and erasing it for Llama-3.3-70B, and a probe-space layer sweep bounds the attainable drift reduction at 6.7--8.9%. We further propose a Moral Representation Consistency (MRC) metric whose underlying framework attributions are validated by human annotators (mean cosine similarity = 0.859), and we report what an automated coherence rater does and does not establish about it.
Figures & tables
| Free Framework | Fixed Framework | |
|---|---|---|
| Structured | A (pilot study) | B (structured + fixed) |
| Unstructured | C (free-form) | D (free-form + fixed) |
| Free Framework | Fixed Framework | |
|---|---|---|
| Structured | 60.8% (A) | 53.4% (B) |
| Unstructured | 53.8% (C) | 54.0% (D) |
| Step | GPT-5 | Llama-3.3-70B | Qwen2.5-72B |
|---|---|---|---|
| 1 | Deont. (27.3) | Virtue (24.5) | Deont. (23.9) |
| 2 | Deont. (23.5) | Virtue (27.4) | Virtue (26.4) |
| 3 | Util. (28.2) | Util. (28.4) | Util. (28.7) |
| 4 | Deont. (28.2) | Virtue (25.1) | Util. (24.6) |
| Model | N | FDR | Entropy | Faithfulness |
|---|---|---|---|---|
| GPT-5 | 1,199 | 0.577 | 1.517 | 0.188 |
| Llama-3.3-70B | 1,200 | 0.567 | 1.500 | 0.191 |
| Qwen2.5-72B | 1,197 | 0.554 | 1.509 | 0.185 |
| Model | Stable Acc | Unstable Acc | Gap |
|---|---|---|---|
| GPT-5 | 73.6% | 66.9% | +6.7 pp |
| Qwen2.5-72B | 54.6% | 52.8% | +1.8 pp |
| Llama-3.3-70B | 62.3% | 64.4% | pp |
| Overall | 63.8% | 61.8% | +2.0 pp |
Appendix figures & tables44 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Moral St. | ETHICS | Social Ch. |
|---|---|---|---|
| GPT-4o | 100.0 | 76.0 | 31.0 |
| GPT-4o-mini | 100.0 | 67.0 | 31.0 |
| GPT-5 | 100.0 | 82.0 | 30.0 |
| GPT-5-mini | 100.0 | 81.0 | 32.0 |
| o3-mini | 100.0 | 81.0 | 31.0 |
| o4-mini | 100.0 | 84.0 | 30.6 |
| Dataset | GPT-5 | Llama | Qwen | Pooled |
|---|---|---|---|---|
| Moral Stories † | 98.8 | 74.8 | 49.9 | 74.6 |
| ETHICS † | 67.4 | 58.4 | 59.9 | 61.9 |
| Social Chem. 101 † | 51.9 | 56.8 | 50.2 | 53.0 |
| Overall | 72.7 | 63.3 | 53.4 | 63.1 |
| GPT-5 | Llama | Qwen | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Sub-task | Acc | FDR | S4 | Acc | FDR | S4 | Acc | FDR | S4 |
| Deontology | 84.0 | .587 | D 52 | 73.0 | .573 | D 38 | 75.0 | .537 | D 42 |
| Justice | 73.0 | .640 | D 38 | 76.0 | .680 | U 35 | 79.0 | .567 | U 36 |
| Utilitarianism † | 72.7 | .583 | D 46 | 60.6 | .537 | U 52 | 57.1 | .543 | U 50 |
| Virtue ‡ | 40.0 | .633 | V 38 | 24.0 | .507 | V 67 | 28.3 | .545 | V 52 |
| Model | Steps | Tokens/Step | Total Length |
|---|---|---|---|
| GPT-5 | 5.14 | 379 | 2,303 |
| GPT-5-mini | 4.60 | 357 | 2,783 |
| GPT-4o | 3.97 | 102 | 2,017 |
| GPT-4o-mini | 4.11 | 109 | 2,115 |
| o4-mini | 4.21 | 156 | 1,848 |
| o3-mini | 3.59 | 168 | 2,127 |
| Dataset | Steps | GPT-4o | GPT-4o-mini | GPT-5 | GPT-5-mini | o3-mini | o4-mini | Total |
|---|---|---|---|---|---|---|---|---|
| ETHICS | 3 | 7 | 0 | 0 | 0 | 41 | 1 | 49 |
| 4 | 87 | 78 | 10 | 38 | 57 | 81 | 351 | |
| 5 | 6 | 22 | 57 | 62 | 2 | 18 | 167 | |
| 6 | 0 | 0 | 25 | 0 | 0 | 0 | 25 | |
| 7 | 0 | 0 | 8 | 0 | 0 | 0 | 8 | |
| Subtotal | 100 | 100 | 100 | 100 | 100 | 100 | 600 |
| Comparison | Diff (pp) | 95% CI | -value | Sig? | ||
| Overall | 618 | 806 | +2.0 | [ 2.8, +6.7] | 0.481 | No |
| GPT-5 only | 212 | 290 | +6.7 | [ 0.8, +14.1] | 0.079 | No |
| Qwen + ETHICS | 72 | 90 | +12.8 | [ 1.5, +26.8] | 0.080 | No |
| By Dataset | ||||||
| ETHICS | 191 | 285 | +2.3 | [ 5.5, +10.2] | 0.555 | No |
| Moral Stories | 243 | 241 | 0.6 | [ 7.5, +6.1] | 0.865 | No |
| Framework | GPT-5 | Llama | Qwen |
|---|---|---|---|
| Deontology | 143 (67.1) | 71 (33.5) | 72 (36.7) |
| Act Utilit. | 42 (19.7) | 73 (34.4) | 82 (41.8) |
| Virtue Ethics | 18 (8.5) | 60 (28.3) | 32 (16.3) |
| Contractualism | 10 (4.7) | 8 (3.8) | 10 (5.1) |
| Contractarian. | 0 (0.0) | 0 (0.0) | 0 (0.0) |
| Total | 213 | 212 | 196 |
| Dataset | GPT-5 | Llama | Qwen |
|---|---|---|---|
| ETHICS | 13.5 | 16.5 | 18.0 |
| Moral Stories | 24.0 | 21.2 | 16.1 |
| Social Chemistry 101 | 15.8 | 15.2 | 15.0 |
| Model | Step | Conseq. | Deont. | Virtue | Care | Social | Dominant |
|---|---|---|---|---|---|---|---|
| GPT-5 | 1 | 55.1 | 68.7 | 53.0 | 52.8 | 50.0 | Deont. |
| GPT-5 | 2 | 52.8 | 52.2 | 55.5 | 49.9 | 40.3 | Virtue |
| GPT-5 | 3 | 76.2 | 68.1 | 58.7 | 51.9 | 50.8 | Conseq. |
| GPT-5 | 4 | 58.0 | 65.4 | 49.6 | 43.8 | 39.3 | Deont. |
| Llama-3.3-70B | 1 | 46.0 | 56.6 | 52.6 | 49.3 | 42.8 | Deont. |
| Llama-3.3-70B | 2 | 41.1 | 39.8 | 60.1 | 51.0 | 34.8 | Virtue |
| Step | GPT-5 | Llama | Qwen |
|---|---|---|---|
| 1 | Deont. (68.7) | Deont. (56.6) | Deont. (50.9) |
| 2 | Virtue (55.5) | Virtue (60.1) | Virtue (52.9) |
| 3 | Conseq. (76.2) | Conseq. (67.2) | Conseq. (60.6) |
| 4 | Deont. (65.4) | Conseq. (57.3) | Conseq. (56.3) |
| Framework | Compliance | Score | Step Compl. | FDR |
|---|---|---|---|---|
| Utilitarianism | 100.0% | 76.5 | 95.8% | 0.108 |
| Deontology | 99.6% | 60.5 | 97.5% | 0.063 |
| Virtue Ethics | 100.0% | 62.5 | 97.3% | 0.083 |
| Contractualism | 94.6% | 49.2 | 82.7% | 0.351 |
| Contractarianism | 94.2% | 53.0 | 81.4% | 0.433 |
| Model | Best Layer | KL | Top-1 | |
|---|---|---|---|---|
| Llama | 63/81 (78%) | 0.123 | 0.527 | 0.457 |
| Qwen | 17/81 (21%) | 0.137 | 0.517 | 0.420 |
| Step 1 | Step 2 | Step 3 | Step 4 | |
| KL Divergence | ||||
| Llama | 0.127 | 0.138 | 0.103 | 0.125 |
| Qwen | 0.094 | 0.171 | 0.149 | 0.135 |
| Top-1 Accuracy | ||||
| Llama | 0.427 | 0.480 | 0.573 | 0.627 |
| Qwen | 0.427 | 0.427 | 0.547 | 0.667 |
| Llama | Qwen | |||
|---|---|---|---|---|
| Category | KL | Top-1 | KL | Top-1 |
| Single-Framework | 0.121 | 0.621 | 0.128 | 0.680 |
| Funnel-to-Util | 0.119 | 0.528 | 0.101 | 0.575 |
| Bounce | 0.115 | 0.474 | 0.134 | 0.417 |
| High-Entropy | 0.087 | 0.429 | 0.134 | 0.250 |
| Other | 0.200 | 0.321 | 0.259 | 0.250 |
| Direction | Transfer KL | Within KL | Degrad. |
|---|---|---|---|
| Llama Qwen | 0.182 | 0.137 | 32.5% |
| Qwen Llama | 0.182 | 0.123 | 47.6% |
| FDR Value | Llama | Qwen |
|---|---|---|
| 0 (consistent) | 212 (17.7%) | 196 (16.4%) |
| 0.33 | 206 (17.2%) | 261 (21.8%) |
| 0.67 | 512 (42.7%) | 493 (41.2%) |
| 1 | 270 (22.5%) | 247 (20.6%) |
| Model | Stable Acc | Unstable Acc | Gap | Gap | Interpretation | |
|---|---|---|---|---|---|---|
| Llama | 0 (baseline) | 62.3% | 61.0% | pp | – | Consistent better |
| 4.0 | 63.0% | 67.1% | pp | pp | Gap reverses | |
| 10.0 | 64.7% | 64.9% | pp | pp | Gap erased | |
| Qwen | 0 (baseline) | 54.6% | 51.8% | pp | – | Consistent better |
| 4.0 | 64.8% | 61.3% | pp | pp | Both arms improve | |
| 5.0 | 62.2% | 58.3% | pp | pp | Gap widens |
| Metric | Value |
|---|---|
| Framework-consistent flip rate | 68.3% (41 of 60 pairs) |
| Drifting flip rate | 88.3% (53 of 60 pairs) |
| Susceptibility ratio | 1.29 |
| Chi-square statistic | 5.94 ( ) |
| Cohen’s (effect size) | 0.498 (small) |
| Metric | Correlation | -value | Interpretation |
|---|---|---|---|
| MRC Score (composite) | Tracks the rater almost exactly, but the rater was given the rule MRC implements | ||
| Stability component | Framework consistency, the component the rater was instructed on | ||
| Drift component (1-FDR) | Transition count, also directly visible to the rater | ||
| Variance component (1-entropy) | Entropy was not shown to the rater, and agreement drops accordingly |
| Category | MRC (mean std) | |
|---|---|---|
| Overall | 3,596 | |
| Single-framework | 621 | |
| Bounce | 2,168 | |
| High-entropy | 807 |
| Model | Dataset | Stable | Unstable | Diff |
|---|---|---|---|---|
| Qwen | ETHICS | 63.9% | 51.1% | +12.8 |
| GPT-5 | Moral St. | 99.0% | 98.7% | +0.2 |
| GPT-5 | Social Ch. | 47.6% | 47.7% | 0.0 |
| Llama | ETHICS | 60.6% | 62.6% | 2.0 |
| Qwen | Social Ch. | 50.0% | 52.1% | 2.1 |
| Llama | Social Ch. | 54.1% | 57.0% | 2.9 |
| Model | Stable | Unstable | Diff | Direction |
|---|---|---|---|---|
| GPT-5 | 78.4% | 61.0% | +17.5 | Expected |
| Qwen | 32.1% | 30.0% | +2.1 | Expected |
| Llama | 45.0% | 55.5% | 10.5 | Opposite |
| Overall | 54.8% | 49.6% | +5.1 | – |
| Framework | Score |
| Kantian Deontology | |
| Benthamite Act Utilitarianism | |
| Aristotelian Virtue Ethics | |
| Scanlonian Contractualism | |
| Gauthierian Contractarianism | |
| Total | 100 |
| Metric | Score |
|---|---|
| Justified | ___ (true/false) |
| Confidence | ___ (0–100) |