Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.
Figures & tables
Figure 1: Controlled study workflow. The curriculum builder creates parallel complexity, task, and prompt views that converge in a matched scheduler. The scheduler instantiates model-family and adaptation choices, launches distributed training or inference, and routes behavioral and selected-task diagnostic evidence to the four research questions. The dashed path denotes the curriculum constraints applied directly to scheduling.
Figure 2: Representative ARC-TGI transformation family. Grid-level variables change object configuration, positions, and colors. Task-level variables determine their joining direction and alignment within each generated episode.
Figure 3: RQ1: hyperparameter sensitivity and leading configurations. Top: reported importance scores for the decoder-only, mixture-of-experts, and encoder–decoder model families. Bottom: the ten highest-accuracy configurations in each sweep. Axes are scaled independently by family, and importance scores are descriptive rather than an additive variance partition. LR, WU, and WD denote learning rate, warm-up ratio, and weight decay.
Figure 4: RQ1: adaptation and acquisition dynamics. (a) Exact-match accuracy for the selected full fine-tuning and LoRA configurations over model-family-specific normalized training progress. Colors identify model families; solid circles denote full fine-tuning and dashed squares denote LoRA. (b) Median generator-level accuracy with interquartile bands over normalized training progress. Endpoint, rank, and compute details appear in Appendix B ; disaggregated generator trajectories appear in Appendix C .
Figure 5: RQ2: family-level heterogeneity. Generators are sorted by exact-match accuracy for each model family. The dashed line is global mean accuracy and dotted lines mark mean ±0.1 . Low dispersion can reflect consistent success or uniform failure.
RQ
Comparison
Condition
D
MoE
ED
RQ1
Training breadth
60 generators
39.3
39.1
7.5
193 generators
51.6
45.0
13.0
Adaptation
Full fine-tuning
53.9
32.6
21.1
LoRA
54.4
32.8
10.8
RQ2
Cross-benchmark transfer
Seen-family val. max.
54.4
32.8
21.1
Observed ARC-AGI peak
3.5
2.75
2.00
Table 1: Behavioral results summary. Exact-match accuracy (%) under the two conditions stated for each RQ1–RQ3 comparison. Complete trajectories, intermediate scales, formulation variants, and compute accounting appear in the corresponding figures and appendices.
Figure 6: RQ3: generator-level effects of in-context experience. Pairwise generator accuracies for one to four demonstrations. Points above or below the diagonal indicate which model family performs better; gray denotes a tie.
Figure 7: RQ4: exploratory layer-wise attention diagnostics. Top: mean attention entropy on the selected small-grid task (upper row) and large-grid task (lower row) after small-grid or large-grid training. Bottom: the baseline-normalized change in mean per-token target-answer log probability after blocking attention from the system message, demonstrations, or test-grid description. Negative knockout values indicate supportive context. Colors identify D (blue), MoE (green), and ED (orange); knockout panels retain native scales.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: RQ1: hyperparameter response distributions. Exact-match accuracy across the tested epoch, learning-rate, warm-up, and weight-decay values for D (top), MoE (middle), and ED (bottom). Points are individual sweep configurations, violins summarize their marginal distributions, and red lines connect marginal means. Each panel marginalizes over the other hyperparameters and should therefore be read as a diagnostic response profile, not an isolated treatment effect.
Family
Update
Epoch
Acc.
Time
GPU
VRAM
Min/epoch
Adj. GPU h
Acc./h
(%)
(sec)
(GB)
D
Full
20
53.90
63303
4 H100
94
52.75
70.34
0.766
LoRA
25
54.36
80191
4 H100
94
53.46
89.10
0.611
MoE
Full
5
32.57
111511
1 H200
141
371.70
61.95
0.526
LoRA
5
32.89
95580
1 H200
141
318.60
53.10
0.618
ED
Full
30
21.12
25365
4 H100
94
14.09
28.18
0.749
Appendix
Table 2: RQ1: unified adaptation and compute summary. Selected full fine-tuning and LoRA endpoints for each model family. Accuracy is exact grid match; time and peak VRAM describe the realized run. Adjusted GPU hours use the stated H200 =2× H100 accounting convention. Acc./h is reported accuracy divided by adjusted GPU hours and is descriptive rather than a hardware-independent efficiency measure.
Figure 9: RQ1: generator-level checkpoint dynamics. Held-out exact-match trajectories for D (blue), MoE (green), and ED (orange) across 20 checkpoints. The six largest improvements are marked with solid circles, and three representative declines or zero-change cases with dashed diamonds; remaining generator trajectories are gray. Counts below each panel summarize improved, decreased, and zero-performing families between the sampled endpoints.
Figure 10: RQ1: experience depth. Generator-level exact-match accuracy across 191 fixed transformation families as training examples per family increase from 10 to 100. Panels correspond to D (blue), MoE (green), and ED (orange). Columns are generators, rows are sample counts, and intensity encodes accuracy on a shared 0–1 scale. Persistent vertical structure shows that generator identity remains influential as within-family experience increases.
Figure 11: RQ1: experience breadth. Generator-level exact-match accuracy on 50 fixed evaluation families as the training curriculum expands from 60 to 193 generators. From left to right, panels show D, MoE, and ED. Columns are curriculum sizes and rows are shared evaluation generators, sorted independently by each model’s 60-generator baseline. Color encodes accuracy on a common 0–1 scale; annotations summarize retained, improved, and markedly decreased endpoint cases.
Generators
D
MoE
ED
60
39.3
39.1
7.5
80
44.8
39.2
8.7
100
43.8
41.5
9.7
140
49.3
43.2
12.7
180
51.0
45.1
12.2
193
51.6
45.0
13.0
Appendix
Table 3: RQ1: aggregate experience-breadth results. Exact-match accuracy (%) on the same 50 evaluation families as the training curriculum expands from 60 to 193 generators. These averages summarize the same conditions shown at generator level in Figure 11 .
Epoch
D
MoE
ED
0
0.5
0.25
0.25
1
2.0
2.75
1.0
3
3.0
1.25
1.75
5
0.5
2.75
1.25
7
–
1.50
–
10
0.25
–
1.50
Appendix
Table 4: RQ2: unified transfer results. Exact-match accuracy (%). Left: original ARC-AGI evaluation at each available checkpoint; dashes denote unmeasured checkpoints. Right: matched and crossed grid-scale evaluation with generator identity shared across conditions. The left panel is a cross-benchmark comparison, whereas the right panel isolates a spatial shift within the filtered ARC-TGI families.
Intervention
Condition
D
MoE
ED
Formulation
Transduction: few-shot
44.3
33.7
0.0
Transduction: reasoning
42.7
32.6
15.6
Induction: few-shot
52.3
59.0
16.1
Induction: reasoning
42.4
59.2
12.6
Context
One shot
36.7
20.1
43.6
Two shots
45.8
26.5
37.9
Appendix
Table 5: RQ3: formulation and in-context experience. Exact-match accuracy (%) under direct grid transduction and executable-rule induction with few-shot or reasoning-description prompts, followed by accuracy on fixed task and generator support as demonstrations increase from one to four.
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China. · Institute of Artificial Intelligence, Xiamen University