cs.CLSep 29, 2026

Large-scale factor analysis shows machine intelligence is only partially interpretable

Authors: Faiz Ghifari Haznitrama, Afrizal Hasbi Azizy, Faeyza Rishad Ardi

Organizations: School of Computing, KAIST · Independent Researcher

Abstract

A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods. A robust pattern across different modes of bias is that 1. A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions, 2. Content-similar benchmarks do not necessarily cluster together, and 3. The gg factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard "intelligence" benchmarks. Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.

Figures & tables

Appendix figures & tables32 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DEPART: DEcomposing PARiTy across Multilingual LLMs

    May 27, 2026Manan Uppadhyay, Prashant Kodali, Pranjal Chitale +3

  2. AwarenessBench: Assessing Cognitive Capabilities of Language Models

    Sep 28, 2026Xiaojian Li, Rongwu Xu, Tianyun Zhang +9Cognitive ScienceMetacognition

  3. ArchitectureIQ: On the Measure of Training Intuition

    Sep 30, 2026Zirui Ren, Shaoyang Guo, Chencheng Tang +9IntuitionArtificial Intelligence Scientists