cs.LGSep 28, 2026

ScAn-Bench: Evaluating Scaling Analysis Methodology

Authors: Artin Sermaxhaj, Nastaran Alipour, Donat Sinani, Johannes Hog, Neeratyoy Mallik, Jenia Jitsev, Danny Stoll

Organizations: University of Freiburg · Zuse School ELIZA · Juelich Supercomputing Center (JSC), Research Center Juelich (FZJ)

Abstract

Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future research, we introduce the surrogate benchmarks ScAn-Bench-LLM and ScAn-Bench-VLM based on 4524 and 8024 checkpoints of language and vision-language model pipelines. On our benchmarks, we perform the first systematic evaluation of both data acquisition and extrapolation methodology for scaling analysis across different data modalities.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

    May 29, 2026Sang Truong, Yuheng Tu, Rylan Schaeffer +1Item Response TheoryScaling Laws

  2. On Test-Time Scaling for Vision-Language Models

    Jun 27, 2026Fawaz Sammani, Tzoulio Chamiti, Nikos DeligiannisTest-Time ScalingRecent Vision-Language Models

  3. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

    Feb 18, 2026Mubashara Akhtar, Anka Reuel, Prajna Soni +34Artificial Intelligence BenchmarksLarge Language Model Benchmarks