How Language Models Organize and Structure Moral Knowledge
Organizations: Distiller Labs
Abstract
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 10 distinct splits exist, so it cannot reject at the 0.05 level) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.
Figures & tables
| construction | mean cosine | PC1 (obs / pred) | peak acc |
|---|---|---|---|
| isotropic floor | (chance) | — | |
| matched non-moral battery | 0.013 [0.005, 0.020] | 0.183 / 0.179 | 1.00 |
| six moral foundations | 0.24 [0.22, 0.25] | 0.388 / 0.387 | 0.94–1.00 |
| shared-neutral-pool (non-moral) | 0.53 [0.51, 0.55] | 0.612 / 0.609 | 1.00 |
| Probe (1B dense) | Raw | RMS-normalized |
|---|---|---|
| Single-foundation | 5.02 | 10.0 |
| Pooled dilemma | 3.81 | 10.0 |
| Per-type dilemma | 3.55 | 9.31 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Layer | Care | Fairness | Liberty | Loyalty | Authority | Sanctity |
|---|---|---|---|---|---|---|
| 0 | 1.000 | 1.000 | 1.000 | 0.938 | 1.000 | 0.938 |
| 1 | 0.875 | 0.938 | 0.938 | 0.938 | 1.000 | 0.938 |
| 2 | 0.938 | 0.938 | 0.938 | 0.938 | 1.000 | 0.938 |
| 3 | 0.938 | 0.938 | 0.938 | 0.938 | 1.000 | 0.938 |
| 4 | 1.000 | 0.938 | 1.000 | 0.938 | 1.000 | 0.938 |
| 5 | 1.000 | 0.938 | 1.000 | 1.000 | 1.000 | 1.000 |
| Layer | Care | Fairness | Liberty | Loyalty | Authority | Sanctity |
|---|---|---|---|---|---|---|
| 0 | 0.938 | 0.875 | 0.938 | 0.938 | 0.875 | 0.938 |
| 1 | 0.938 | 0.875 | 0.812 | 0.812 | 0.812 | 0.812 |
| 2 | 0.938 | 0.875 | 0.875 | 0.812 | 0.875 | 0.812 |
| 3 | 0.938 | 0.938 | 0.938 | 0.875 | 0.875 | 0.875 |
| 4 | 0.938 | 0.938 | 0.938 | 1.000 | 1.000 | 1.000 |
| 5 | 0.938 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| Layer | Care | Fairness | Liberty | Loyalty | Authority | Sanctity |
|---|---|---|---|---|---|---|
| 0 | 0.740 ∗ | 0.741 ∗ | 0.768 ∗ | 0.761 ∗ | 0.780 ∗ | 0.793 ∗ |
| 1 | 0.752 ∗ | 0.765 ∗ | 0.775 ∗ | 0.763 ∗ | 0.780 ∗ | 0.789 ∗ |
| 2 | 0.774 ∗ | 0.779 ∗ | 0.796 ∗ | 0.785 ∗ | 0.790 ∗ | 0.789 ∗ |
| 3 | 0.767 ∗ | 0.788 ∗ | 0.791 ∗ | 0.784 ∗ | 0.808 | 0.811 |
| 4 | 0.789 ∗ | 0.803 | 0.817 | 0.796 ∗ | 0.812 | 0.819 |
| 5 | 0.799 ∗ | 0.814 | 0.812 | 0.818 | 0.809 | 0.830 |
| Care | Fair | Lib | Loy | Auth | Sanc | |
|---|---|---|---|---|---|---|
| Care | 1.000 | 0.148 | 0.189 | 0.203 | 0.164 | 0.201 |
| Fair | 0.148 | 1.000 | 0.232 | 0.179 | 0.224 | 0.143 |
| Lib | 0.189 | 0.232 | 1.000 | 0.273 | 0.333 | 0.224 |
| Loy | 0.203 | 0.179 | 0.273 | 1.000 | 0.264 | 0.219 |
| Auth | 0.164 | 0.224 | 0.333 | 0.264 | 1.000 | 0.249 |
| Sanc | 0.201 | 0.143 | 0.224 | 0.219 | 0.249 | 1.000 |
| Care | Fair | Lib | Loy | Auth | Sanc | |
|---|---|---|---|---|---|---|
| Care | 1.000 | 0.264 | 0.221 | 0.250 | 0.188 | 0.229 |
| Fair | 0.264 | 1.000 | 0.236 | 0.294 | 0.292 | 0.195 |
| Lib | 0.221 | 0.236 | 1.000 | 0.317 | 0.324 | 0.252 |
| Loy | 0.250 | 0.294 | 0.317 | 1.000 | 0.329 | 0.242 |
| Auth | 0.188 | 0.292 | 0.324 | 0.329 | 1.000 | 0.298 |
| Sanc | 0.229 | 0.195 | 0.252 | 0.242 | 0.298 | 1.000 |
| Care | Fair | Lib | Loy | Auth | Sanc | |
|---|---|---|---|---|---|---|
| Care | 1.000 | 0.246 | 0.200 | 0.226 | 0.168 | 0.195 |
| Fair | 0.246 | 1.000 | 0.220 | 0.347 | 0.278 | 0.138 |
| Lib | 0.200 | 0.220 | 1.000 | 0.352 | 0.393 | 0.237 |
| Loy | 0.226 | 0.347 | 0.352 | 1.000 | 0.415 | 0.220 |
| Auth | 0.168 | 0.278 | 0.393 | 0.415 | 1.000 | 0.295 |
| Sanc | 0.195 | 0.138 | 0.237 | 0.220 | 0.295 | 1.000 |
| Layer | Observed statistic | -value |
|---|---|---|
| 0 | 0.001 | 0.40 |
| 1 | 0.014 | 0.80 |
| 2 | 0.003 | 0.70 |
| 3 | 0.004 | 0.80 |
| 4 | 0.004 | 0.80 |
| 5 | 0.003 | 0.50 |
| Experiment | Model | Wall time |
|---|---|---|
| 1–3 (probing + geometry + bootstrap) | OLMo-2 1B | 5 min |
| 5 (dense vs. MoE geometry) | OLMoE-1B-7B | 20 min |
| 6 (geometric trajectory) | OLMo-2 1B (37 ckpts) | 45 min |
| 7 (framework fragility) | OLMo-2 + OLMoE | 25 min |
| Layer | Matched | Mismatched | Gap |
|---|---|---|---|
| 0 | 0.0664 | 0.0313 | +0.0351 |
| 1 | 0.0723 | 0.0334 | +0.0389 |
| 2 | 0.0861 | 0.0450 | +0.0411 |
| 3 | 0.0939 | 0.0525 | +0.0414 |
| 4 | 0.1007 | 0.0507 | +0.0500 |
| 5 | 0.1057 | 0.0494 | +0.0563 |