Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout
Organizations: University of Southern California · Duke University · University of Michigan
Abstract
A concept subspace's effect on model behavior does not establish how it relates to the output readout. We introduce a two-sided geometric diagnostic that measures an extracted subspace's overlap with the dominant right-singular directions of the unembedding matrix, evaluated against output-oriented positive controls. Given an extracted basis, the raw diagnostic requires only model weights. Our testbed is the Format-Agnostic Reasoning Subspace (FARS), a ten-dimensional basis extracted from eighteen reasoning concepts expressed in six surface forms. Across nine rank-matched estimators and twenty-six models, four activation-derived concept estimators carry only 0.38--0.80% mean energy in the top-ten readout span. Final-layer PCA carries 3.56%, exceeding FARS in 25 of 26 models. A same-layer next-token control, evaluated using a fitted linear translator for depth matching, carries approximately thirteen times more energy than FARS, with separation in all 25 tested models. Re-extracting FARS on ten disjoint concepts yields 62--100% cross-format retrieval across twenty-four generative models, demonstrating transfer of the extraction procedure rather than a fixed basis. A complementary four-model, three-seed intervention study finds model-dependent source-directed effects that remain well below full-vector replacement. Together, the geometry and intervention controls distinguish concept structure from dominant readout directions while limiting claims of causal sufficiency.
Figures & tables
| Model | Argmax preserved (%) | Behavioural (drop pp) | |||
| best FARS layer | FARS-10 | PCA-10 | Full | GSM8K | MATH |
| GPT-2 XL | 93.7 | 87.7 | 81.0 | — | — |
| Qwen-2.5-3B | 89.7 | 37.3 | 22.0 | — | — |
| Qwen-2.5-7B-Inst. | 94.7 | 59.7 | 26.3 | pp | pp |
| Mistral-7B | 97.3 | 85.3 | 83.3 | — | — |
| Llama-3.1-8B | 94.0 | 56.7 | 39.7 | (n.s.); at | — |
| vs. | vs. | ||||
|---|---|---|---|---|---|
| Rank- subspace | role | energy % | energy % | ||
| Final-layer PCA | positive ctrl | 4.556 | 3.56 | 0.86 | 65.5 |
| FormPCA | control | 4.738 | 1.26 | 0.90 | 75.8 |
| FullPCA | control | 4.751 | 1.04 | 0.83 | 77.5 |
| Shuffled-FARS | control | 4.762 | 0.90 | 0.62 | 78.9 |
| Diff-of-means | published | 4.774 | 0.80 | 0.46 | 79.5 |
| Probe | Llama pair | Qwen pair | QwQ-32B | Baseline |
|---|---|---|---|---|
| (base last, distill last) | — | (rank- ceiling) | ||
| (last, CoT) | (rank- ceiling) | |||
| CoT-tail FARS % cross-format Agn. | – (Haar rand.) | |||
| X-FARS pool Llama-Distill | ( pp) | — | (15 models) | |
Appendix figures & tables78 assets
Supplementary material from the paper’s appendix.
Appendix
| Form | Text |
|---|---|
| English | If it rains then the ground is wet. It rains. Therefore the ground is wet. |
| Chinese | 如果下雨,那么地面是湿的。下雨了,所以地面是湿的。 |
| French | S’il pleut alors le sol est mouillé. Il pleut. Donc le sol est mouillé. |
| Python | def mp(p, q): return q if p else None |
| Math | |
| Struct. | P1: P->Q | P2: P | Rule: MP | Q |
| Domain | Concept | Example |
| Arithmetic | Multi-step eval | |
| Modular | ||
| Proportional | 3 cost $12; 7 cost? | |
| GCD | ||
| Logic | Syllogism | All A are B; all C are A |
| Modus ponens | P Q; P; therefore Q |
| Domain | Concept | Example (math form) |
|---|---|---|
| Temporal | Event ordering | , min, |
| Duration arithmetic | ||
| Probabilistic | Conditional probability | |
| Expected value | ||
| Graph | Shortest path | |
| Connectivity | ; |
| Model | Space | Dim | Concept-ARI | Concept% | Form% |
|---|---|---|---|---|---|
| GPT-2 XL | Full | 1600 | .142 | 28.7 | 46.0 |
| FARS | 10 | .300 | 42.6 | 25.9 | |
| Form ctrl | 5 | .038 | 17.6 | 92.6 | |
| Qwen-7B | Full | 3584 | .284 | 42.3 | 25.9 |
| FARS | 10 | .619 | 66.7 | 23.8 | |
| Form ctrl | 5 | .036 | 16.7 | 95.4 |
| Model | RSA-C | Probe% | Purity% | FARS Patch | |
|---|---|---|---|---|---|
| Qwen2.5-7B (base) | .150 | 59.8 | 66.7 | .943 | 28.8 |
| +Instruct | .153 | 62.9 | 68.8 | .867 | 25.5 |
| Llama-3.1-8B (base) | .165 | 53.9 | — | .941 | — |
| +Instruct | .187 | 59.8 | 67.6 | .890 | 12.6 |
| Benchmark | Direction | Full top-1 | FARS top-1 | |
|---|---|---|---|---|
| TriForm in-dist. ( pairs, avg.) | all | |||
| MBPP-sanitized ( ) | prose code | |||
| MBPP-sanitized ( ) | code prose |
| FARS | Control | |||||
|---|---|---|---|---|---|---|
| Model | RSA-C | RSA-F | Probe | Boost | RSA-F | RSA-C |
| GPT-2 XL | .317 | .118 | 52.8 | +20.1 | .626 | .003 |
| Qwen-3B | .357 | .056 | 64.4 | +11.6 | .634 | .005 |
| Qwen-7B | .367 | .009 | 68.0 | +8.2 | .638 | .010 |
| Mistral | .366 | .008 | 70.5 | +10.2 | .631 | .003 |
| Llama-8B | .368 | .015 | 71.9 | +18.0 | .634 | .007 |
| Model | Metric | FARS | Shuffled | LDA | Full-PCA |
|---|---|---|---|---|---|
| GPT-2 XL | RSA-C | .279 | .381 | .091 | |
| RSA-F | .107 | .473 | .011 | .486 | |
| Qwen-7B | RSA-C | .357 | .384 | .151 | |
| RSA-F | .020 | .243 | .012 | .261 | |
| Mistral-7B | RSA-C | .362 | .380 | .194 | |
| RSA-F | .012 | .259 | .012 | .275 |
| Model | Family | FARS energy | |||
|---|---|---|---|---|---|
| Llama-3.1-70B-Inst | dense | 23/80 | |||
| OLMo-2-7B-Inst | dense | 13/32 | |||
| Falcon3-7B-Inst | dense | 8/27 | |||
| Mistral-7B-v0.3 | dense | 11/31 | |||
| Mistral-7B-Inst | dense | 10/32 | |||
| Phi-4 | dense | 12/39 |
| energy % in top- of | energy % in top- of | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | FARS | NextTok ℓ | Last-PCA | FARS | LDA | FullPCA | Haar | NextTok ℓ | ||
| Phi-4 | 12/39 | 0.40 | 0.40 | 1.19 | 4.29 | 0.22 | 0.25 | 0.61 | 0.26 | |
| Llama-3.1-70B-Inst | 22/79 | 0.37 | 0.16 | 0.72 | 1.65 | 0.19 | 0.19 | 0.45 | 0.10 | |
| Mistral-7B-Inst | 9/31 | 0.34 | 0.40 | 6.57 | 4.58 | 0.37 | 0.29 | 1.03 | 0.25 | |
| Falcon-Mamba-7B | 29/63 | 0.50 | 0.28 | 0.45 | 0.31 | 0.32 | 0.28 | 0.61 | 0.22 | |
| Llama-3.1-8B-Inst | 8/31 | 0.44 | 0.49 | 0.84 | 4.51 | 0.29 | 0.24 | 0.52 | 0.30 | |
| enrichment , FARS / final-layer | Haar | ||||
|---|---|---|---|---|---|
| Model | |||||
| DeBERTa-v3-large | 1024 | 0.83 / 1.2 | 0.92 / 1.1 | 0.93 / 1.1 | 1.03 |
| DeepSeek-V2-Lite | 2048 | 1.26 / 11.7 | 1.12 / 6.3 | 1.06 / 3.2 | 1.00 |
| Falcon-Mamba-7B | 4096 | 1.17 / 16.8 | 1.36 / 6.8 | 1.23 / 3.4 | 1.01 |
| Falcon3-7B-Inst | 3072 | 0.96 / 10.7 | 0.95 / 5.8 | 1.00 / 2.9 | 0.98 |
| gpt-oss-20B | 2880 | 0.77 / 15.7 | 0.88 / 10.3 | 0.83 / 5.4 | 1.00 |
| AUC | Haar | vs. | vs. | final block, vs. | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | ref. | truth | refusal | truth | refusal | truth | NextTok 1 | refusal | truth | ||
| DeepSeek-V2-Lite | 13 | 1.00 | 1.00 | 0.49 | 0.59 | 1.42 | 1.04 | 1.01 | 1.48 | 7.55 | 8.37 |
| Falcon-Mamba-7B | 29 | 1.00 | 1.00 | 0.24 | 0.40 | 0.36 | 2.12 | 0.11 | 2.81 | 1.06 | 5.77 |
| Falcon3-7B-Inst | 8 | 1.00 | 1.00 | 0.33 | 0.41 | 0.59 | 1.06 | 0.18 | 0.02 | 1.33 | 5.54 |
| gpt-oss-20B | 7 | 1.00 | 1.00 | 0.35 | 0.29 | 0.24 | 0.51 | 0.15 | 2.61 | 4.05 | 3.84 |
| GPT-2 XL | 20 | 1.00 | 1.00 | 0.62 | 3.92 | 0.33 | 4.28 | 0.28 | 0.30 | 7.86 | 1.44 |
| Haar | energy % in top- readout span | |||||
|---|---|---|---|---|---|---|
| Model | FARS | random-token | vocabulary-anchored | ratio | ||
| Llama-3.1-70B-Inst | 8192 | 0.12 | 0.16 | 2.37 | 77.3 | |
| SmolLM3-3B | 2048 | 0.49 | 0.64 | 26.49 | 40.7 | |
| R1-Distill-Qwen-14B | 5120 | 0.20 | 0.23 | 1.99 | 27.0 | |
| Llama-3.1-8B-Inst | 4096 | 0.24 | 0.49 | 4.27 | 26.4 | |
| gpt-oss-20B | 2880 | 0.35 | 0.27 | 5.99 | 25.9 | |
| Rank- subspace | energy % | ||
|---|---|---|---|
| SAE dictionary (top- PCA of live decoder directions) | 4.590 | 3.93 | 56.3 |
| SAE concept features (selectivity-weighted, rank ) | 4.748 | 1.06 | 73.8 |
| FARS at the same layer | 4.759 | 0.96 | 75.1 |
| SAE random features (control) | 4.749 | 0.88 | 76.5 |
| Haar expectation | — | 0.24 | — |
| Model | RSA-C | RSA-F | Probe-W% | Probe-X% | X/W | Agn% |
|---|---|---|---|---|---|---|
| GPT-2 XL | .100 | .430 | 96.6 | 32.7 | .34 | 20.0 |
| Qwen2.5-3B | .126 | .501 | 98.5 | 52.8 | .54 | 23.3 |
| Qwen2.5-7B | .150 | .485 | 99.1 | 59.8 | .60 | 19.5 |
| Mistral-7B | .200 | .529 | 98.5 | 60.3 | .61 | 20.8 |
| Llama-3.1-8B | .165 | .525 | 99.4 | 53.9 | .54 | 21.1 |
| Model | frac. | Agn% peak | peak | Agn% last | last | |
|---|---|---|---|---|---|---|
| Llama-3.1-8B-Inst | 9/ 32 | 0.29 | 81.2% | 4.79 | 66.4% | 4.63 |
| Mistral-7B | 9/ 32 | 0.29 | 81.8% | 4.81 | 69.4% | 4.63 |
| Llama-3.1-70B-Inst | 24/ 81 | 0.30 | 87.3% | 4.87 | 65.4% | 4.74 |
| Mistral-7B-Inst | 11/ 33 | 0.34 | 84.0% | 4.81 | 61.1% | 4.63 |
| Falcon3-7B-Inst | 10/ 28 | 0.37 | 79.0% | 4.82 | 32.7% | 4.65 |
| Mixtral-8x7B-Inst | 12/ 33 | 0.38 | 82.4% | 4.73 | 30.2% | 4.73 |
| Rank | Variance (% of =17) | Agn% (% of =17) |
|---|---|---|
| 52.9 | 52.7 | |
| 71.2 | 60.5 | |
| 84.3 | 73.4 | |
| 89.9 | 83.0 | |
| 94.2 | 89.4 | |
| 100.0 | 100.0 |
| Model | Variance (% of =17) | Agn% (% of =17) |
|---|---|---|
| qwen2.5-7b | 89.1 | 82.5 |
| qwen2.5-3b-instruct | 91.7 | 74.9 |
| phi-3.5-mini-instruct | 91.3 | 86.0 |
| mistral-7b-v0.3 | 87.9 | 86.1 |
| gpt2-xl | 91.3 | 85.8 |
| olmo-2-7b-instruct | 87.6 | 88.4 |
| Model | Consec in-plat | – in-plat | Outside | Random |
|---|---|---|---|---|
| GPT-2 XL | ||||
| Qwen-3B-Inst | ||||
| Phi-3.5-mini-Inst | ||||
| Mamba-2.8B | ||||
| Mistral-7B-base | ||||
| Mistral-7B-Inst |
| Model | next token | language id | sentiment | factual tru | inference v | mono. | |
|---|---|---|---|---|---|---|---|
| gpt2-xl | 1600 | 14.4 | 8.2 | 4.0 | 1.4 | 1.0 | ✓ |
| phi-3.5-mini-instruct | 3072 | 13.5 | 35.4 | 18.0 | 9.7 | 7.2 | — |
| qwen2.5-7b | 3584 | 62.9 | 2.1 | 4.8 | 33.9 | 38.0 | — |
| llama-3.1-8b-instruct | 4096 | 17.3 | 2.1 | 4.6 | 3.8 | 1.5 | — |
| mistral-7b-v0.3 | 4096 | 19.3 | 37.5 | 26.1 | 55.3 | 15.0 | — |
| olmo-2-7b-instruct | 4096 | 27.4 | 4.0 | 2.1 | 14.2 | 6.3 | — |
| Model | FARS | Random | Norm-matched | Full vector | No intervention |
|---|---|---|---|---|---|
| GPT-2 XL | 10.8 [7.5, 16.7] | 0.0 | 0.0 | 41.9 [34.2, 47.5] | 0.0 |
| Qwen2.5-3B-Inst. | 3.3 [0.8, 7.5] | 0.0 | 0.0 | 47.8 [43.3, 51.7] | 0.0 |
| SmolLM3-3B | 0.8 [0.0, 1.7] | 0.0 | 0.0 | 19.7 [16.7, 22.5] | 0.0 |
| Mamba-2.8B | 0.8 [0.0, 1.7] | 0.0 | 0.0 | 26.9 [21.7, 33.3] | 0.0 |
| Top-1% | Top-10 ov. | KL med. | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | FARS | PCA | Full | FARS | PCA | Full | FARS | PCA | Full |
| GPT-2 XL | 93.7 | 87.7 | 81.0 | .90 | .65 | .47 | .016 | .154 | .655 |
| Qwen-3B | 89.7 | 37.3 | 22.0 | .92 | .60 | .44 | .012 | .612 | 2.77 |
| Qwen-7B | 94.7 | 59.7 | 26.3 | .94 | .63 | .48 | .006 | .262 | .683 |
| Mistral-7B | 97.3 | 85.3 | 83.3 | .96 | .74 | .56 | .002 | .034 | .142 |
| Llama-8B | 94.0 | 56.7 | 39.7 | .94 | .64 | .47 | .004 | .098 | .811 |
| Source-directed events (%) | Median | ||||
| Model | Pairs | FARS | Haar | FARS | Haar |
| GPT-2 XL | 40 | 5.0 | 0.0 | 0.02435 | 0.00128 |
| Model | Baseline | FARS abl. | Random abl. | FARS |
|---|---|---|---|---|
| GPT-2 XL | 29.4 | 12.0 | 29.6 | 17.4 |
| Qwen2.5-7B | 52.8 | 18.3 | 52.7 | 34.5 |
| Mistral-7B | 52.8 | 15.6 | 52.9 | 37.2 |
| Llama-3.1-8B-Instruct | 56.5 | 17.7 | 56.4 | 38.8 |
| Ablation | Top-1 | Top-10 overlap | Median KL |
|---|---|---|---|
| Random 10-dim (control, mean of 5 draws) | 0.988 | 0.984 | 0.0005 |
| FARS 10-dim | 0.767 | 0.804 | 0.065 |
| Difference | 22 pp | 18 pp |
| Model | Base. | Rand. abl. | FARS abl. | Paired test | |
|---|---|---|---|---|---|
| Qwen2.5-7B-Instruct ( ) | 100 | 90.0 | 90.0 | 70.0 | |
| Llama-3.1-8B-Instruct ( ) | 100 | 86.0 | 85.5 | 85.0 | (n.s.) |
| Model (layer, ) | Coef. | FARS | Random | Paired | |
| Qwen-7B-Inst. ( , ) | .700 | .900 | |||
| .880 | .910 | ||||
| .870 | .910 | ||||
| .710 | .900 | ||||
| Llama-8B-Inst. ( , ) | .850 | .855 | |||
| .850 | .865 |
| Pair | Jaccard | Patching | Class |
|---|---|---|---|
| en math | .228 | .696 | decl |
| en zh | .108 | .319 | decl |
| en code | .216 | .195 | proc |
| code math | .259 | .208 | proc |
| Jaccard does not predict patching (Pearson , n.s.); code math has | |||
| the highest Jaccard yet second-lowest patching, opposite to a tokenisation confound. | |||
| Comparison | Mean overlap | Paired test |
|---|---|---|
| EN math_notation | .646 | — |
| EN py_decl | .194 | — |
| EN py_code | .140 | — |
| py_decl py_code (procedural axis) | , CI | |
| math_notation py_decl (surface axis) | ||
| math_notation py_code (total gap) |
| Model | Prose refusal | Code refusal | Discordant (p/c) | (prose code) | |
|---|---|---|---|---|---|
| Qwen2.5-7B-Instruct | 100 | .980 | .970 | 1 / 0 | .500 (n.s.) |
| Llama-3.1-8B-Instruct | 200 | .935 | .965 | 1 / 7 | .996 (reject) |
| Quantity | Llama-3.1-8B-Inst. ( ) | Mistral-7B-v0.3 ( ) |
|---|---|---|
| Layer FARS-fraction | ||
| MLP share of layer FARS | ||
| Attention share of layer FARS | ||
| Top-3 head share | ||
| Top-5 head share | ||
| Top-3 head indices | H1, H14, H8 | H31, H4, H0 |
| Intervention | Accuracy | vs. baseline |
|---|---|---|
| Baseline (no hook) | — | |
| Full FARS ablation ( -d) | ||
| Canonical axis ( -d) | ||
| Canonical axis ( -d, “procedural”) | ||
| Canonical axis ( -d) | ||
| Canonical axis ( -d) |
| Model | Peak / Writer | Plateau |
|---|---|---|
| GPT-2 XL | 20 / 18 | 16–25 |
| Qwen-2.5-3B-Inst | 22 / 22 | — |
| Phi-3.5-mini-Inst | 18 / 17 | — |
| Qwen-2.5-7B | 13 / 12 | 9–19 |
| Mistral-7B-v0.3 | 11 / 10 | 6–19 |
| Mistral-7B-Inst | 10 / 8 | — |
| Model | Layer | Frozen FARS | Full space | Random mean |
|---|---|---|---|---|
| GPT-2 XL | 20 | 22.8 | 48.3 | 20.2 |
| Qwen2.5-7B | 13 | 40.0 | 90.6 | 35.3 |
| Mistral-7B-v0.3 | 11 | 39.4 | 86.7 | 30.3 |
| Model | Family | TriForm-FARS | Novelty-FARS | Random | ||
|---|---|---|---|---|---|---|
| Agn% | X/W | Agn% | X/W | Agn% | ||
| R1-Distill-Llama-8B | reasoning | 79.6 | 0.77 | 50.6 | ||
| Mixtral-8x7B-Inst | MoE | 82.1 | 0.81 | 56.1 | ||
| Llama-3.1-70B-Inst | dense | 82.1 | 0.78 | 56.1 | ||
| SmolLM3-3B | dense | 82.4 | 0.80 | 54.4 | ||
| Yi-1.5-9B-Chat | dense | 75.0 | 0.72 | 45.0 | ||
| trunc. | collapse | trace | cross-format (%) | Haar | ||
|---|---|---|---|---|---|---|
| Model | % | % | forms | by prompt | by trace | % |
| R1-Distill-Llama-8B | 79 | 56 | 6 | 30.6 | 26.9 | 8.0 |
| R1-Distill-Qwen-7B | 86 | 57 | 6 | 44.4 | 28.1 | 5.6 |
| Pair | Base layer | Distilled layer | |
|---|---|---|---|
| Llama-3.1-8B / R1-Distill-Llama-8B | 22 | 20 | 3.06 |
| Qwen2.5-7B / R1-Distill-Qwen-7B | 7 | 13 | 3.48 |
| Model | Last-input layer | CoT-tail layer | |
|---|---|---|---|
| R1-Distill-Llama-8B | 20 | 14 | 4.53 |
| R1-Distill-Qwen-7B | 13 | 10 | 4.26 |
| Model | Last-input Agn. (%) | CoT-tail Agn. (%) | Haar (%) |
|---|---|---|---|
| R1-Distill-Llama-8B | 79.6 | 60.8 | 21.6 |
| R1-Distill-Qwen-7B | 75.0 | 58.0 | 24.4 |
| Pool | X-FARS (%) | Haar mean SD | Haar p95 | Gap (pp) |
|---|---|---|---|---|
| 15 models | 53.83 | 27.58 | 28.01 | |
| 16 (+Llama-Distill) | 55.57 | 28.01 | 29.63 | |
| 17 (+Qwen-Distill) | 51.83 | 26.22 | 27.19 |
| Model | Stimuli | Mean length | Median length | Budget hit (%) |
|---|---|---|---|---|
| R1-Distill-Llama-8B | 324 | 489 | 512 | 84.9 |
| R1-Distill-Qwen-14B | 324 | 389 | 512 | 63.9 |
| R1-Distill-Qwen-7B | 324 | 491 | 512 | 88.0 |
| Model | Dim full | Full top-1 | FARS top-1 | |
| GPT-2 XL | 1600 | .351 | .562 | |
| Qwen2.5-7B | 3584 | .495 | .729 | |
| Mistral-7B-v0.3 | 4096 | .522 | .743 | |
| Llama-3.1-8B-Instruct | 4096 | .559 | .738 | |
| Llama-3.1-70B-Instruct | 8192 | .615 | .816 | |
| Pooled ( pairs) | — | .508 | .717 |
| Quantity | Real | Null (shuffled, ) |
|---|---|---|
| FARS GPA, models ( – B) | ( CI ) | |
| FARS GPA, models, arch. families | ||
| FARS GPA, all models | ||
| X-FARS joint, all models | — | |
| FARS GPA, all models | — | |
| X-FARS joint, all models | — |
| Ax. | Var | extreme | extreme |
|---|---|---|---|
| 0 | set_diff, set_ , func_comp | causal_intervention, spatial_containment | |
| 1 | func_comp, multi_step, proportional | set_diff, set_ , causal_confound | |
| 2 | causal_intervention, spatial_containment | logic_negation, arith_modular | |
| 3 | causal_intervention, arith_modular | spatial_direction, spatial_rotation | |
| 4 | modus_ponens, contrapositive, syllogism | causal_confound, logic_negation |
| Objective | Initialization | Final loss | Max. gradient |
|---|---|---|---|
| Raw centroids | Historical | 5.1923 | |
| Raw centroids | Cached FARS | 798.21 | |
| Raw centroids | Random | 69.723 | |
| Unit-norm centroids | Historical | 1.4303 | |
| Unit-norm centroids | Cached FARS | 1.4503 | |
| Unit-norm centroids | Random | 1.4274 |
| Query Corpus | GPT-2 XL | Qwen-7B | Mistral-7B | Llama-8B | Mixtral-MoE | Llama-70B |
|---|---|---|---|---|---|---|
| GPT-2 XL | , | |||||
| Qwen2.5-7B | , | |||||
| Mistral-7B-v0.3 | , | |||||
| Llama-3.1-8B-Inst | , | |||||
| Mixtral-8x7B-Inst (MoE) | , | |||||
| Llama-3.1-70B-Inst | , |
| Metric | FARS + GPA | X-FARS |
|---|---|---|
| Canonical variance explained ( models, arch. families) | ||
| Cross-family pooled (pred. patching, ) | ||
| per-model concept-RSA vs. FARS | , | |
| Steps required | PCA + iterated Procrustes | iterated alignment heuristic |
| Model | Span(A,B) | Span(rand C) | Span(rand dirs) | Both |
|---|---|---|---|---|
| Original -pair set ( stimuli): | ||||
| GPT-2 XL | ( ) | |||
| Llama-3.1-8B-Inst | ( ) | |||
| Mistral-7B-v0.3 | ( ) | |||
| Extended -pair set ( stimuli): | ||||
| Llama-3.1-8B-Inst | ( ) | |||
| Model | Span(A,B,C) | rand-3-C | rand-3-D | vs rand-D | All |
|---|---|---|---|---|---|
| Phi-3.5-mini-Inst (3.8B dense, ) | ( ) | ||||
| Qwen-2.5-3B-Inst (dense, ) | ( ) | ||||
| Mamba-2.8B (state-space, ) | ( ) | ||||
| Qwen-7B (dense decoder, ) | ( ) | ||||
| Mistral-7B-Inst (dense, ) | ( ) | ||||
| Llama-3.1-8B-Inst (dense, ) | ( ) |
| Model | Span(A,B,C,D) | rand-C | rand-D | rand-C | All 4 |
|---|---|---|---|---|---|
| Mamba-2.8B (state-space) | ( ) | ( ) | |||
| Qwen-7B (dense) | (n.s.) | ( ) | |||
| Mistral-7B-Inst (dense) | (n.s.) | ( ) | |||
| Llama-3.1-8B-Inst (dense) | (n.s.) | ( ) |
| / | GPT-2 | Llama | Mistral | Qwen |
|---|---|---|---|---|
| GPT-2 XL | ||||
| Llama-Inst | ||||
| Mistral-7B | ||||
| Qwen-7B | ||||
| Bold: cross-family . Plain: . | ||||
| Projection | Cross-family | |
|---|---|---|
| FARS centroid (10-dim, supervised) | ||
| Random orthonormal (10-dim, no labels) | ||
| Shuffled- (1000 permutations) | (mean) | 95% CI |