How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 10 distinct splits exist, so it cannot reject at the 0.05 level) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.
Figures & tables
Figure 1: Pairwise cosine similarity between foundation probe directions at layer 7 (bootstrap-stable; all six foundations exceed the 0.8 stability threshold at this layer; Section 4.6 ), OLMo-2 1B. Mean off-diagonal cosine = 0.262. See Appendix C for matrices at layers 0 and 15.
construction
mean cosine
PC1 (obs / pred)
peak acc
isotropic floor
∼0
≈1/6 (chance)
—
matched non-moral battery
0.013 [0.005, 0.020]
0.183 / 0.179
1.00
six moral foundations
0.24 [0.22, 0.25]
0.388 / 0.387
0.94–1.00
shared-neutral-pool (non-moral)
0.53 [0.51, 0.55]
0.612 / 0.609
1.00
Table 1: Calibration ladder for the mean pairwise cosine (OLMo-2 1B, layer 7, probe-weight directions, six concepts per construction; paired bootstrap over the 32 direction-estimation pairs, n=200 ). The moral shared component sits ∼20× above the matched non-moral null and below the shared-neutral-pool construction. Mean-cosine entries are bootstrap means; the direct layer-7 moral point is 0.26 (Figure 1 ), the ∼0.02 gap to the bootstrap mean being resampling attenuation, which widens rather than narrows the moral − non-moral difference. PC1 columns give observed vs. the closed-form [1+(k−1)cˉ]/k .
Figure 2: Layer-wise geometric metrics for OLMo-2 1B. (a) Mean pairwise cosine similarity is relatively flat across layers. (b) Effective dimensionality remains constant at 5 across all layers. (c) MFT group structure: mean cosine within the individualizing group, within the binding group, and between groups track together across layers, so the directions do not separate into the predicted individualizing/binding clusters.
Figure 3: Hierarchical clustering (Ward’s method) of foundation probe directions at layer 7 (bootstrap-stable), OLMo-2 1B. The first split does not recover the MFT individualizing/binding distinction.
Figure 4: Mean critical noise σ∗ per foundation for OLMo-2 1B (left) and OLMoE-1B-7B (right), seed-averaged with cap-at-max and bootstrap 95% CIs over noise seeds. Every foundation is more fragile in MoE than in dense (the cross-architecture dilution effect); within each architecture the per-foundation CIs overlap and the ordering is not statistically separable.
Figure 5: Geometric trajectory during OLMo-2 1B pre-training (37 checkpoints, seed-fixed foundation directions). (a) Mean pairwise cosine reaches the integration regime by step 2–5K, then slowly differentiates. (b) Effective dimensionality remains at 5 throughout. (c) Accuracy keeps climbing after the geometry sets.
Figure 6: Subspace membership scores for 15 dilemma probe directions across 16 layers. Each cell shows the fraction of a dilemma direction’s variance explained by the 2D subspace of its component foundation directions. Liberty–sanctity shows the strongest compositional signal.
Figure 7: Mean dilemma subspace membership across layers: matched (component foundations, red, ±1 SD band), the mismatched-pair baseline (blue dashed), and the random-vector null (gray, ∼0.001 ). Matched membership exceeds the mismatched baseline at every layer; both sit far above the random null.
Figure 8: Component balance at each pair’s peak subspace membership layer. Values near 0.5 (dashed line) indicate balanced contribution from both foundations. Fairness–sanctity (red) is the only pair with substantial imbalance.
Figure 9: Distribution of pairwise cosine similarities between dilemma directions at layer 13, split by whether the pairs share a component foundation. Shared-component pairs (blue, n=60 ) have higher mean similarity (0.273) than non-sharing pairs (red, n=45 , mean 0.196); exact permutation p=0.0001 .
Figure 10: Hierarchical clustering of all 21 directions (6 foundation in blue, 15 dilemma in red) at layer 13. Foundation directions cluster separately from dilemma directions.
Figure 11: Raw mean critical noise for probes at three complexity levels (single-foundation from Exp. 7, pooled and per-type dilemma). The apparent gradient does not survive RMS normalization (see text): under scale correction the single-foundation and pooled-dilemma values converge, indicating the raw ordering reflects register-dependent activation scale rather than encoding robustness.
Probe (1B dense)
Raw σ∗
RMS-normalized σ∗
Single-foundation
5.02
10.0
Pooled dilemma
3.81
10.0
Per-type dilemma
3.55
9.31
Table 13
Figure 12: Foundation direction cosine similarity for OLMo-2 1B (layer 7/16) and 7B (layer 14/32), at matched relative depth. Mean off-diagonal cosine is 0.262 (1B) and 0.240 (7B) at these display layers; both sit in the integration range. Effective dimensionality is 5 for both (Figure 13 ).
Figure 13: Layer-wise geometry across scale on normalized depth. (a) Mean pairwise cosine for both models stays in the integration band. (b) Effective dimensionality is constant at 5 for both 1B and 7B.
Figure 14: Per-foundation mean critical noise σ∗ for OLMo-2 1B and 7B (dense), seed-averaged with bootstrap 95% CIs. The 7B model is uniformly more robust (scale effect); within each model the per-foundation CIs overlap and no foundation is reliably most robust.
Figure 15: Discovered clusters (rows) against MFT foundations (columns) at the most foundation-aligned layer (11) of OLMo-2 7B. Cells are pair counts; row-normalized shading. The clusters do not map onto foundations (AMI = 0.032).
Figure 16: MFV external replication on OLMo-2 7B. (a) Cosine similarity among the six MFV foundation directions. (b) Mean cosine and effective dimensionality vs depth for MFV and for matched mean-difference directions on our dataset. (c) Per-foundation alignment between MFV and our directions.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Layer
Care
Fairness
Liberty
Loyalty
Authority
Sanctity
0
1.000
1.000
1.000
0.938
1.000
0.938
1
0.875
0.938
0.938
0.938
1.000
0.938
2
0.938
0.938
0.938
0.938
1.000
0.938
3
0.938
0.938
0.938
0.938
1.000
0.938
4
1.000
0.938
1.000
0.938
1.000
0.938
5
1.000
0.938
1.000
1.000
1.000
1.000
Appendix
Table 3: Per-foundation probe accuracy across layers, OLMo-2 1B. Each cell shows test accuracy (16 test examples per foundation). All foundations exceed 0.6 at every layer.
Layer
Care
Fairness
Liberty
Loyalty
Authority
Sanctity
0
0.938
0.875
0.938
0.938
0.875
0.938
1
0.938
0.875
0.812
0.812
0.812
0.812
2
0.938
0.875
0.875
0.812
0.875
0.812
3
0.938
0.938
0.938
0.875
0.875
0.875
4
0.938
0.938
0.938
1.000
1.000
1.000
5
0.938
1.000
1.000
1.000
1.000
1.000
Appendix
Table 4: Per-foundation probe accuracy across layers, OLMoE-1B-7B.
Layer
Care
Fairness
Liberty
Loyalty
Authority
Sanctity
0
0.740 ∗
0.741 ∗
0.768 ∗
0.761 ∗
0.780 ∗
0.793 ∗
1
0.752 ∗
0.765 ∗
0.775 ∗
0.763 ∗
0.780 ∗
0.789 ∗
2
0.774 ∗
0.779 ∗
0.796 ∗
0.785 ∗
0.790 ∗
0.789 ∗
3
0.767 ∗
0.788 ∗
0.791 ∗
0.784 ∗
0.808
0.811
4
0.789 ∗
0.803
0.817
0.796 ∗
0.812
0.819
5
0.799 ∗
0.814
0.812
0.818
0.809
0.830
Appendix
Table 5: Bootstrap direction stability (mean cosine similarity with full-data direction, 200 resamples). Values marked with ∗ fall below the 0.8 stability threshold.
Care
Fair
Lib
Loy
Auth
Sanc
Care
1.000
0.148
0.189
0.203
0.164
0.201
Fair
0.148
1.000
0.232
0.179
0.224
0.143
Lib
0.189
0.232
1.000
0.273
0.333
0.224
Loy
0.203
0.179
0.273
1.000
0.264
0.219
Auth
0.164
0.224
0.333
0.264
1.000
0.249
Sanc
0.201
0.143
0.224
0.219
0.249
1.000
Appendix
Table 6: Cosine similarity matrix at layer 0 (peak separation), OLMo-2 1B. Mean off-diagonal = 0.216.
Care
Fair
Lib
Loy
Auth
Sanc
Care
1.000
0.264
0.221
0.250
0.188
0.229
Fair
0.264
1.000
0.236
0.294
0.292
0.195
Lib
0.221
0.236
1.000
0.317
0.324
0.252
Loy
0.250
0.294
0.317
1.000
0.329
0.242
Auth
0.188
0.292
0.324
0.329
1.000
0.298
Sanc
0.229
0.195
0.252
0.242
0.298
1.000
Appendix
Table 7: Cosine similarity matrix at layer 7, OLMo-2 1B. Mean off-diagonal = 0.262.
Care
Fair
Lib
Loy
Auth
Sanc
Care
1.000
0.246
0.200
0.226
0.168
0.195
Fair
0.246
1.000
0.220
0.347
0.278
0.138
Lib
0.200
0.220
1.000
0.352
0.393
0.237
Loy
0.226
0.347
0.352
1.000
0.415
0.220
Auth
0.168
0.278
0.393
0.415
1.000
0.295
Sanc
0.195
0.138
0.237
0.220
0.295
1.000
Appendix
Table 8: Cosine similarity matrix at layer 15, OLMo-2 1B. Mean off-diagonal = 0.262.
Layer
Observed statistic
p -value
0
0.001
0.40
1
− 0.014
0.80
2
− 0.003
0.70
3
− 0.004
0.80
4
− 0.004
0.80
5
0.003
0.50
Appendix
Table 9: Exact permutation test for individualizing/binding group structure across layers, OLMo-2 1B (enumeration over all 20 assignments; attainable floor p=0.10 ). No layer reaches significance.
Experiment
Model
Wall time
1–3 (probing + geometry + bootstrap)
OLMo-2 1B
∼ 5 min
5 (dense vs. MoE geometry)
OLMoE-1B-7B
∼ 20 min
6 (geometric trajectory)
OLMo-2 1B (37 ckpts)
∼ 45 min
7 (framework fragility)
OLMo-2 + OLMoE
∼ 25 min
Appendix
Table 26
Layer
Matched
Mismatched
Gap
0
0.0664
0.0313
+0.0351
1
0.0723
0.0334
+0.0389
2
0.0861
0.0450
+0.0411
3
0.0939
0.0525
+0.0414
4
0.1007
0.0507
+0.0500
5
0.1057
0.0494
+0.0563
Appendix
Table 10: Per-layer matched vs. mismatched dilemma subspace membership, OLMo-2 1B (mean over 15 dilemma pairs). Matched directions are seed-averaged probe-weight directions; mismatched is averaged over the foundation pairs sharing no component.