Autoregressive Large Language Models (LLMs) frequently struggle with deterministic multi-step algorithmic tasks such as multi-digit multiplication and long division. In this paper, we investigate the mechanics of multi-step arithmetic in compact "Tiny" Transformers (~10.6M non-embedding parameters, 49.3M total) trained on synthetic data across four basic operations (+, -, *, /) unrolled as step-by-step scratchpads. First, we establish the necessary training foundations: (1) dataloader sequence padding creates an 83% gradient starvation artifact that collapses accuracy from 40% to 1%, remediated via continuous sequence packing; (2) linguistic pretraining is an essential prerequisite (<= 2.0% without it); and (3) modern architectural primitives (RoPE, RMSNorm, SwiGLU) and Sparse Mixture of Experts (MoE) substantially improve additive reasoning over baseline GPT-2. Second, we demonstrate that algorithmic scratchpad formulation directly dictates success. Introducing a deterministic Digit-by-Digit Long Division scratchpad within a 4-stage Hierarchical Developmental Curriculum dramatically elevates single-digit division from 4.0% to 86.7% accuracy on a 4,000-problem held-out benchmark. In contrast, multi-digit multiplication remained challenging: detailed error analysis revealed that while the model correctly computed single-digit sub-products and place-value zeros, our FOIL scratchpad failed because it forced a simultaneous summation of up to nine multi-digit terms in a single step without pairwise intermediate accumulation. Finally, we identify two key boundaries: performance collapses to 0.00% on unseen 4-digit operands, and unbuffered training induces catastrophic forgetting, collapsing division accuracy from 86.7% down to 0.00%.
Figures & tables
Figure 1 : Comparison of data loading strategies for 256-token context blocks.
Metric
Continuous Streaming
Single-Problem Padded
Overall Accuracy
40.0%
1.0%
Addition Accuracy
84.0%
4.0%
Subtraction Accuracy
76.0%
0.0%
Active Math Tokens / Step
1,024
∼ 180
Total Active Tokens (5k steps)
5,120,000
∼ 900,000 (5.7 × drop)
Gradient Sparsity
0%
82.8%
Table 1: Empirical collapse under single-instruction padded training.
Figure 2 : Active token throughput in millions (blue, left axis) versus downstream arithmetic accuracy (red, right axis) across dataloader packaging strategies over 5,000 steps. Padded training induces severe gradient starvation, collapsing performance.
Run
Architecture
Reasoning Dataset
Wiki Steps
Math Steps
Add (%)
Sub (%)
Mul (%)
Div (%)
Overall (%)
1
Baseline GPT-2
Standard FOIL
20,000
5,000
8.0
12.0
0.0
0.0
5.0
2
FlashAttention
Standard FOIL
20,000
5,000
4.0
20.0
0.0
0.0
6.0
3
Modern Dense
Standard FOIL
20,000
5,000
56.0
76.0
0.0
0.0
33.0
4
MoE (Top-2)
Standard FOIL
20,000
5,000
60.0
84.0
0.0
0.0
36.0
5
Modern Dense
Extended FOIL
20,000
100,000
100.0
100.0
8.0
8.0
54.0
6
MoE (Top-2)
Extended FOIL
20,000
100,000
100.0
100.0
8.0
8.0
54.0
Table 2: Master experimental ablation results across all 10 evaluation runs on the 100-problem evaluation set. Runs 1–4 evaluate baseline scratchpad configurations; Runs 9 and 10 introduce our deterministic digit-by-digit long division algorithm, elevating division from 4.0% to 60.0%.
WikiText Pretraining Steps
Baseline GPT-2
FlashAttention
Modern Dense
Sparse MoE (Top-2)
10 steps (Cold Start)
0.0%
1.0%
2.0%
2.0%
5,000 steps
2.0%
2.0%
32.0%
—
10,000 steps
2.0%
1.0%
31.0%
37.0%
20,000 steps (Optimal Scaffold)
5.0%
6.0%
33.0%
36.0%
50,000 steps (Linguistic Saturation)
—
2.0%
19.0%
33.0%
Table 3: Impact of natural language pretraining duration on downstream arithmetic accuracy after 5,000 mathematical fine-tuning steps. Insufficient pretraining ( ≤10 steps) results in near-total failure across all architectures ( ≤2.0% ), while 20,000 steps provides the optimal syntactic foundation.
Architecture Configuration
Addition (%)
Subtraction (%)
Multiplication (%)
Division (%)
Overall (%)
Run 1: Baseline GPT-2
8.0
12.0
0.0
0.0
5.0
Run 2: + FlashAttention
4.0
20.0
0.0
0.0
6.0
Run 3: + RoPE + RMSNorm + SwiGLU
56.0
76.0
0.0
0.0
33.0
Run 4: + Sparse MoE (Top-2 Routing)
60.0
84.0
0.0
0.0
36.0
Table 4: Architectural ablation at constant ∼ 10.6M non-embedding parameter scale. Modern inductive biases (RoPE, RMSNorm, SwiGLU) and Sparse MoE provide a substantial boost over standard GPT-2.
Figure 3 : Left: Overall multi-task accuracy plateau from 10,000 to 100,000 steps across Dense and MoE architectures, illustrating the sharp 55% ceiling. Right: Per-operation accuracy trajectory in Run 8 (MoE), demonstrating that Addition and Subtraction saturate at 100% while Multiplication remains at 0% and Division caps at 20%.
Figure 4 : Learning trajectories across mathematical operations throughout the 4-stage Hierarchical Developmental Curriculum for Modern Dense (left, Run 9) and Sparse MoE (right, Run 10). Notice the steep rise of Long Division (red triangles) to 56–60% as the training corpus transitions into Stage 4.
Figure 5 : Long Division Performance: Comparison of Division accuracy across algorithmic paradigms in ∼ 10M parameter networks. Transforming division from partial quotient guessing to deterministic digit-by-digit descent produces a substantial improvement (0–4% → 60.0% on development tests).
Task / Operational Slice
Evaluated Samples
Modern Dense (Run 9)
Sparse MoE (Run 10)
Step 30,000 (Peak)
Step 50,000 (Final)
Step 30,000 (Peak)
Step 50,000 (Final)
Addition ( ≤3 digits, In-Dist)
514
93.39% (480/514)
99.22% (510/514)
91.44% (470/514)
98.83% (508/514)
Addition (4 digits, Out-of-Dist)
486
0.00% (0/486)
0.00% (0/486)
0.00% (0/486)
0.00% (0/486)
Subtraction ( ≤3 digits, In-Dist)
517
98.07% (507/517)
99.42% (514/517)
96.52% (499/517)
99.42% (514/517)
Subtraction (4 digits, Out-of-Dist)
483
1.86% (9/483)
1.66% (8/483)
0.62% (3/483)
1.66% (8/483)
Multiplication (1–3 digits, All)
1,000
11.20% (112/1,000)
2.50% (25/1,000)
5.10% (51/1,000)
2.50% (25/1,000)
Table 5: Large-scale 4,000-problem benchmark comparing Step 30,000 (peak multi-task) and Step 50,000 (final) checkpoints across Modern Dense and Sparse MoE architectures. At Step 30,000, single-digit long division achieves up to 86.67% accuracy. Prolonged training to Step 50,000 drives addition and subtraction toward 99.4% saturation but induces catastrophic developmental drift on division. Unseen 4-digit addition completely collapses (0.00%), defining a strict length generalization boundary.
Figure 6 : Large-Scale Benchmark Dynamics across 4,000 held-out problems. (a) Length Generalization Collapse: While models achieve near-deterministic accuracy ( >99.2% ) on in-distribution operands ( ≤3 digits), accuracy abruptly collapses to 0.00% on unseen 4-digit operands due to fixed positional and causal depth horizons. (b) Catastrophic Developmental Drift: Single-digit division performance peaks at Step 30,000 (86.67% Dense, 84.33% MoE), but prolonged un-buffered training to Step 50,000 reduces division accuracy down to 0.00% as late-stage addition gradients overwrite parameters.