cs.LGOct 6, 2026

A Systematic Study of Small Language Models on Abstract Reasoning Tasks

Authors: Nur A Zarin Nishat, Jens Lehmann, Andrei Aioanei, Sahar Vahdati

Organizations: Leibniz University of Hanover TIB – Leibniz Information Centre for Science and Technology and University Library · Amazon

Abstract

Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning

    May 29, 2026Saku Peltonen, August Bøgh Rønberg, Andreas Plesner +1Multiplex Graph TransformersGraph Representations

  2. DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

    Jun 25, 2026Yuxuan Yang, Feiyang Li, Yile WangReasoning SkillsNegative Results

  3. A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation

    May 17, 2026Qingchuan Ma, Yuexiao Ma, Yongkang Xie +3Reasoning BenchmarkLarge Language Model Benchmarks