Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future research, we introduce the surrogate benchmarks ScAn-Bench-LLM and ScAn-Bench-VLM based on 4524 and 8024 checkpoints of language and vision-language model pipelines. On our benchmarks, we perform the first systematic evaluation of both data acquisition and extrapolation methodology for scaling analysis across different data modalities.
Figures & tables
Model family
Surrogate query (CPU-s)
Training (GPU-h)
VLM
29
419.58
LLM
17
3519.42
Table 1 : Runtime comparison for 100 evaluations. Surrogate querying is measured in CPU hours on AMD EPYC 9655 CPU using 8 cores, while training cost is reported in GPU hours on A100 GPUs.
Benchmark
#HPs
#Scale
FLOP Range
Cont.
Disc.
#Down.
ScAn-Bench-VLM
6
3
3.6×1013 – 7.8×1018
7
2
40
ScAn-Bench-LLM
5
4
3.2×1016 – 2.9×1020
6
3
-
Table 2 : Summary of the two surrogate benchmarks. The table reports the number of hyperparameters (#HPs), scale parameters (#Scale), FLOP ranges, continuous and discrete dimensions, and the number of downstream tasks (#Down.) modeled by each surrogate benchmark.
VLM
LLM
Surrogate
Spearman ↑
RMSE ↓
Spearman ↑
RMSE ↓
TabPFN
0.98
0.21
0.95
0.27
AutoGluon
0.98
0.26
0.91
0.30
Ensemble (XGB)
0.96
0.35
0.85
0.35
Ensemble (Mix)
0.95
0.37
0.86
0.33
Ensemble (LGB)
0.96
0.42
0.87
0.33
Table 3: Comparison of surrogate performance across model families in predicting upstream performance, measured by Spearman rank correlation (higher is better) and RMSE (lower is better).
LLM
OpenCLIP
Strategy
0.1x
0.5x
Best
0.1x
0.5x
Best
CARBS
2.02±0.23
1.76±0.17
1.31
2.92±0.41
2.49±0.14
0.41
RS
1.71±0.18
1.89±0.37
1.31
2.28±0.29
1.81±0.30
0.41
Table 4 : Actual loss at predicted configurations. Evaluation of full-parameter extrapolations using CARBS and Random Search (RS) at various Ctarget budget ratios.
Table 10 : Fixed training choices across all LLM configurations.
Training
Downstream
Upstream
Total
GPU
VLM
LLM
VLM
LLM
VLM
LLM
RTX 2080 Ti
11.84
-
731.84
-
11.39
-
755.07
RTX 3080
142.42
-
373.53
-
13.47
-
529.42
L40
0.44
-
-
-
9.84
-
10.28
A100 40GB
3986.02
11042.45
828.75
935.35
198.66
3972.34
20962.55
H100 94GB
6653.11
26870.62
859.65
-
159.61
348.27
34891.26
Appendix
Table 11 : Compute cost (GPU hours) for data collection across different GPU types. The total column reports the sum of training and evaluation costs across both VLM and LLM benchmarks.
Family
Surrogate
RMSE ↓
MAE ↓
MDAE ↓
MARPD ↓
R2↑
R ↑
Corr. ↑
VLM
TabPFN
0.21
0.06
0.02
2.69
0.96
0.98
0.98
AutoGluon
0.26
0.12
0.07
5.79
0.94
0.97
0.98
XGB
0.35
0.20
0.13
9.11
0.90
0.95
0.96
Mix
0.37
0.21
0.14
10.11
0.89
0.95
0.95
LGB
0.42
0.29
0.24
14.88
0.86
0.95
0.96
LLM
TabPFN
0.27
0.08
0.01
3.02
0.77
0.88
0.95
Appendix
Table 12: Surrogate performance across model families on held-out final configurations.
Category
Field
Description
Input
Configuration
Sampled from ScAn space A.1
Training progress
Fidelity in [0,1] indicating training progress
Output
Upstream metrics
Validation or test loss
Downstream metrics
Defined in A.4
FLOPs
Cumulative training compute up to queried fidelity
Model parameters
Number of model parameters
Appendix
Table 13 : Surrogate API query interface. Given a configuration and a fidelity value, the API returns performance predictions and additional metadata.