Large-scale factor analysis shows machine intelligence is only partially interpretable
Organizations: School of Computing, KAIST · Independent Researcher
Abstract
A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods. A robust pattern across different modes of bias is that 1. A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions, 2. Content-similar benchmarks do not necessarily cluster together, and 3. The factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard "intelligence" benchmarks. Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.
Figures & tables
| Dataset | Imputer | AVE | ||||
|---|---|---|---|---|---|---|
| C Std. | Mean fill | 14 | 5.5% | 0.708 | 0.142 | 0.224 |
| C Std. | missForest | 4 | 22.1% | 0.695 | 0.398 | 0.399 |
| S Std. | SoftImpute (corr.) | 5 | 9.3% | 0.676 | 0.306 | 0.378 |
| C Std. | Zero fill | 14 | 5.4% | 0.621 | 0.093 | 0.286 |
| C Aggr. | missForest | 4 | 19.3% | 0.521 | 0.112 | 0.241 |
| C Std. | k-NN | 7 | 10.8% | 0.516 | 0.209 | 0.288 |
| Label | Median A | Significant |
|---|---|---|
| code | +0.264 | 6/18 |
| math | +0.031 | 0/18 |
| reasoning | -0.008 | 0/18 |
| No | Benchmark | 95% CI | Best | Worst | |||
|---|---|---|---|---|---|---|---|
| 1 | bhasa | -0.346 | 0.141 | [-0.005, 0.287] | 0.024 | 0.318 | 5 |
| 2 | mtrag | -0.341 | 0.147 | [-0.051, 0.346] | 0.021 | 0.394 | 5 |
| 3 | creativityprism | -0.339 | 0.156 | [-0.019, 0.332] | 0.026 | 0.367 | 5 |
| 4 | eqbench | -0.332 | 0.156 | [-0.060, 0.372] | 0.017 | 0.451 | 5 |
| 5 | mceval | -0.331 | 0.156 | [0.024, 0.288] | 0.051 | 0.333 | 5 |
| 6 | pwc_svamp | -0.329 | 0.169 | [-0.033, 0.371] | 0.058 | 0.298 | 4 |
Appendix figures & tables32 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Correlation evidence | Decision | Rows removed |
|---|---|---|---|
| LiveCodeBench release windows v1–v6 (Kaggle) | mean pairwise , worst pair , over 45 shared models | Keep the aggregate, drop 6 per-version identifiers | 270 |
| TwitterAAE dialect splits (African-American and White English) | – with each other and the parent, over 32 shared models | Keep the parent, drop both dialect splits | 64 |
| GPQA variants (few/zero-shot diamond/main, Kaggle) | mean , worst pair , over 46–47 shared models, with the better-populated canonical GPQA and GPQA Diamond already present | Drop all 4 Kaggle variants | 185 |
| MultiLoKo per-language splits (31 languages, Kaggle) | mean pairwise , near-duplicate for well-resourced pairs (Simplified/Traditional Mandarin , Italian/Swedish ) and noisy for low-resource pairs on small overlap | Keep the paper-sourced MultiLoKo across-language aggregate, drop all 31 per-language identifiers | 1,523 |
| A Kaggle re-import of MMLU | Cross-source re-import (44 rows) of the canonical MMLU column (468 rows) | Drop the re-import | 44 |
| A Kaggle re-import of SciCode | Not a third metric. Per-model value matching shows it splices main-problem scores for 30 of 46 models and sub-problem scores for the other 13, which is a scraping artifact | Drop as a data-integrity fix. The 4 explicit split variants ( – ) are kept , as their correlations are not uniform enough to treat as duplicates | 46 |
| Method | Description | Package | Configuration |
|---|---|---|---|
| SoftImpute ( Mazumder et al., 2010 ) | Nuclear-norm-penalized low-rank completion by iterative soft-thresholded SVD, assuming a low-rank signal plus noise. Primary cell-level method. | softImpute | sweeps rank: 1…10 (capped at ), and at each rank a 30-point geometric grid from down to , ALS with warm starts |
| k-NN | Each missing cell filled from the most similar models, an assumption-light baseline with no low-rank, linearity, or normality assumption. | VIM | sweeps : 1…10 (capped below ), Gower distance over benchmarks, weighted-mean aggregation |
| missForest ( Stekhoven & Bühlmann, 2012 ) | Iterative random-forest imputation, nonparametric, able to capture nonlinear dependence the low-rank methods cannot represent. | missForest | sweeps number of trees: {50, 100, 200, 400}, at most 10 iterations |
| Estimator | Description | Package | Configuration |
|---|---|---|---|
| One-sided matrix completion ( Cao et al., 2023 ) | Recovers the right-singular vectors of a reduced matrix of a large, super-sparse dataset. Originally demonstrated to work with simply 2 observations per row. | Custom Julia code | Sweeps rank 1…10, selecting the rank with the best . |
| SoftImpute ( Mazumder et al., 2010 ) | Applies SoftImpute’s low-rank completion to the observed pairwise correlation matrix rather than the data matrix, whose missing entries are exactly the benchmark pairs never co-observed. | softImpute | sweeps rank 1…10 with the same nested grid as the cell-level methods above |
| USVT ( Chatterjee, 2015 ) | Universal singular value thresholding: completes the correlation matrix by hard-thresholding its singular values. | filling | fixed singular-value threshold , no sweep |
| Densifier | Benchmarks | 2023 | 2024 | ||
|---|---|---|---|---|---|
| C | 78 | 87/20%/10 | 196/14%/6 | 294/12%/21 | 93/12%/10 |
| S | 124 | 86/14%/42 | 196/9%/30 | 292/9%/26 | 94/10%/16 |
| Densifier | Imputer | 2023 | 2024 | ||
|---|---|---|---|---|---|
| C | SoftImpute | -1.33 | -0.29 | 0.59 | 0.00 |
| C | missForest | -1.02 | -0.40 | 0.22 | 1.11 |
| C | k-NN | -1.36 | -0.34 | 0.25 | 1.22 |
| S | SoftImpute | -1.47 | -0.21 | 0.37 | 0.67 |
| S | missForest | 0.43 | -0.32 | -0.19 | 0.86 |
| S | k-NN | -0.52 | 0.06 | -0.27 | 1.20 |
| Densifier | Imputer | Era | Year gap ( ) | Co-obs. | |
|---|---|---|---|---|---|
| C | SoftImpute | -0.16 | -6.1 | 0.07 (0.022) | 0.39 |
| C | missForest | 0.21 | -4.2 | 0.10 (0.054) | 0.31 |
| C | k-NN | -0.34 | -8.2 | 0.06 (0.082) | 0.43 |
| C | OneSidedMC | 0.10 | -12.6 | 0.15 (0.000) | 0.54 |
| C | SoftImpute (corr.) | -0.06 | -6.1 | 0.11 (0.018) | 0.25 |
| S | SoftImpute | -0.03 | -6.9 | 0.09 (0.004) | 0.20 |
| Method | Dataset | |||
|---|---|---|---|---|
| Mean fill | C Std. | 2 | +0.4192 | 78 |
| Mean fill | C Std. | 14 | +0.4192 | 78 |
| k-NN | C Std. | 2 | +0.4065 | 78 |
| k-NN | C Std. | 7 | +0.4065 | 78 |
| k-NN | S Std. | 2 | +0.2335 | 124 |
| k-NN | S Std. | 11 | +0.2335 | 124 |
| Method | Dataset | ||
|---|---|---|---|
| Mean fill | C Std. | 78 | -0.429 |
| k-NN | C Std. | 78 | -0.429 |
| k-NN | S Std. | 124 | +0.020 |
| missForest | C Aggr. | 102 | -0.342 |
| missForest | C Std. | 78 | -0.106 |
| missForest | S Std. | 124 | +0.261 |
| cells | benchmarks | raw | adjusted |
|---|---|---|---|
| 2 | 54 | -0.023 | -0.021 |
| 3 | 26 | -0.024 | -0.049 |
| 4 | 24 | -0.214 | -0.232 |
| 5 | 141 | +0.061 | +0.046 |
| 8 | 10 | +0.040 | +0.038 |
| 10 | 33 | +0.518 | +0.516 |
| Label | Median A | Significant | Median n |
|---|---|---|---|
| conversation | +0.283 | 3/6 | 7 |
| likelihood_probe | +0.153 | 6/10 | 19 |
| short_qa | +0.098 | 2/18 | 15 |
| classification | +0.095 | 2/18 | 15 |
| interactive | +0.078 | 1/10 | 16 |
| long_reasoning | +0.066 | 0/18 | 9 |
| Label | Median A | Significant | Median n |
|---|---|---|---|
| monolingual_non_english | +0.291 | 16/18 | 19 |
| multilingual | +0.204 | 6/14 | 9 |
| crosslingual | +0.137 | 1/10 | 12 |
| english | -0.034 | 0/18 | 85 |
| Label | Median A | Significant | Median n |
|---|---|---|---|
| finance | +0.870 | 10/10 | 4 |
| professional_writing | +0.632 | 4/6 | 5 |
| fact_verification | +0.300 | 0/4 | 5 |
| commonsense | +0.287 | 5/18 | 11 |
| code | +0.264 | 6/18 | 11 |
| translation | +0.249 | 6/10 | 16 |
| Dataset | Imputer | RMSE | Configuration | |
|---|---|---|---|---|
| S Std. | SoftImpute | 0.6338 | 0.504 | rank=5 (swept 1..10) |
| C Std. | SoftImpute | 0.6575 | 0.493 | rank=9 (swept 1..10) |
| S Std. | missForest | 0.6798 | 0.471 | ntree=400 (swept [50,100,200,400]) |
| C Std. | missForest | 0.7321 | 0.399 | ntree=50 (swept [50,100,200,400]) |
| S Std. | SoftImpute (corr.) | 0.7637 | 0.378 | rank=6 (swept 1..10) |
| S Std. | OneSidedMC | 0.7554 | 0.365 | r=2 (swept 1..10) |
| Method | Dataset | |||
|---|---|---|---|---|
| Mean fill | C Std. | 2 | +0.0701 | 78 |
| Mean fill | C Std. | 14 | +0.1316 | 78 |
| k-NN | C Std. | 2 | +0.0268 | 78 |
| k-NN | C Std. | 7 | +0.3205 | 78 |
| k-NN | S Std. | 2 | +0.1765 | 124 |
| k-NN | S Std. | 11 | +0.2069 | 124 |
| No | Benchmark | 95% CI | Best | Worst | Reference | |||
|---|---|---|---|---|---|---|---|---|
| 1 | bhasa | -.346 | .141 | [-.005, .287] | .024 | .318 | 5 | Leong et al.,2023 |
| 2 | mtrag | -.341 | .147 | [-.051, .346] | .021 | .394 | 5 | Katsis et al.,2025 |
| 3 | creativityprism | -.339 | .156 | [-.019, .332] | .026 | .367 | 5 | Hou et al.,2026 |
| 4 | eqbench | -.332 | .156 | [-.060, .372] | .017 | .451 | 5 | Paech,2023 |
| 5 | mceval | -.331 | .156 | [.024, .288] | .051 | .333 | 5 | Chai et al.,2025 |
| 6 | pwc_svamp | -.329 | .169 | [-.033, .371] | .058 | .298 | 4 | Patel et al.,2021 |