We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.
Figures & tables
Figure 1: Capability and inference efficiency of compact language models. Left: scores across eight evaluation categories. Upper right: decoding throughput relative to IronLLM-0.6B at 32K context. Lower right: benchmark-level score differences versus relative generation time, compared with IronLLM-0.6B MTP.
Figure 2: Overview of the IronLLM-0.6B architecture. Gated DeltaNet (GDN) linear-attention layers and gated global attention (GA) layers are interleaved at a 3:1 ratio, with Zero-RMSNorm adopted throughout for training stability. An X-MTP module with a shared attention block and tied embedding/LM-head weights provides multi-token prediction.
Component
IronLLM-0.6B
Total parameters
654M
Number of layers
24
Hidden size
1024
Intermediate size
3584
Attention type
Hybrid
Positional encoding
Partial RoPE
Table 1: Core architecture configuration of IronLLM-0.6B.
Figure 3: Overview of the IronLLM-0.6B-Light architecture. Building on IronLLM-0.6B, the Light variant replaces all RMSNorm layers with Dynamic Tanh (DyT), removes the attention output gate and QK-Norm, omits the SiLU activation after the causal convolution in GDN, and adopts the upper-bounded ReLUx activation, reducing inference cost and improving quantization friendliness.
Figure 4: Training Data Ecosystem.
Figure 5: Illustration of the multi-stage pre-training pipeline. (a) Data composition across stages; (b) corresponding learning-rate schedule following the Warmup-Stable-Decay (WSD) pattern.
Checkpoint
MMLU
CMMLU
CEval
BBH
GSM8K
MATH
MBPP
Step X−10k
52.32
48.39
48.11
35.65
51.55
19.06
36.19
Step X
53.18
48.90
47.67
36.52
50.34
17.64
40.47
Xmerge
54.17
49.93
49.49
37.70
57.85
22.82
40.47
Gain over Step X
+0.99
+1.03
+1.82
+1.18
+7.51
+5.18
+0.00
Table 2: Performance comparison across benchmarks. Best results are highlighted in bold .
Benchmark
# Shots
Mode
IronLLM-0.6B Base
IronLLM-0.6B -Light Base
MiniCPM5-1B Base 1 1 1 The officially released MiniCPM5-1B-Base checkpoint predates mid-training stage.
Qwen3.5-0.8B Base
Qwen3-0.6B Base
General Tasks
MMLU
5-shot
PPL
55.67
50.62
45.21
49.94
54.46
MMLU-Pro (CoT)
5-shot
Gen
26.45
23.24
21.07
25.71
23.93
MMLU-redux
5-shot
PPL
57.11
51.57
46.61
51.02
55.69
ARC-Challenge
0-shot
PPL
64.41
61.02
48.14
72.54
66.10
ARC-Easy
0-shot
PPL
79.72
72.66
62.26
84.83
83.07
Table 3: Comparison of IronLLM-0.6B-Base with other representative base models. Bold and underlined values indicate the best and second-best non-thinking results, respectively.
Stage
Context Window
RoPE Base
Tokens
Pretrain
4,096
10,000
6.2T
Stage 1
32,768
10,000→1,000,000
42B
Stage 2
65,536
1,000,000 (fixed)
21B
Table 4: Context extension schedule. The RoPE base is scaled at Stage 1 and kept fixed at Stage 2.
Figure 6: Overview of the IronLLM post-training pipeline. IronLLM-0.6B follows three stages: General SFT, independent domain-specialist training, and Multi-Domain On-Policy Distillation (MOPD). During MOPD, frozen specialists provide token-level guidance through teacher prefill, while domain verifiers supply sequence-level verifiable rewards (VRs) on student-generated trajectories. IronLLM-0.6B-Light skips specialist training and reuses the specialists trained from IronLLM-0.6B.
Figure 7: Composition of the general supervised fine-tuning (SFT) dataset.
Benchmark (Metric)
IronLLM-0.6B 2 2 2 For the IronLLM models, we use temperature=0.7, top_p=0.8, top_k=-1, presence_penalty=1.5, and repetition_penalty=1.0. All other baseline models use their officially recommended sampling parameters.
IronLLM-0.6B-Light 2
Qwen3-0.6B 3 3 3 Scores of Qwen3, Qwen3.5, and MiniCPM5 are reported as Non-thinking / Thinking.
LFM2-700M
Qwen3.5-0.8B 3 , 4 4 4 We observed relatively low performance on both code and mathematics benchmarks for Qwen3.5, accompanied by severe repetition artifacts in its thinking traces across both domains. The official Qwen3.5 blog reports no code evaluation results and marks mathematics scores as “–”, indicating that the scores are not yet available or not applicable. Our results are consistent with those reported in the MiniCPM5 blog.
MiniCPM5-1B 3
Long Context
RULER 5 5 5 RULER scores are averaged over 4K, 8K, 16K, 32K, and 64K context lengths.
87.7
82.1
40.4 / 60.0
56.2
87.5 / 82.1
67.5 / 78.2
LongBench v2
27.0
27.2
28.4 / 28.2
16.1
27.8 / 25.8
24.3 / 26.4
General Knowledge
MMLU-Pro
42.1
35.9
24.4 / 37.6
21.5
35.4 / 45.9
36.8 / 47.5
MMLU-Redux
60.4
56.2
46.5 / 56.0
47.1
54.2 / 63.3
59.7 / 69.5
Table 5: Comparison of IronLLM-0.6B with representative post-trained models. Bold and underlined values indicate the best and second-best non-thinking results, respectively.
Model
Attention
Decode speed at position L (tokens/s)
TPOT growth 1K → 32K
1K
4K
16K
32K
IronLLM-0.6B
GDN hybrid
460
445
410
377
1.22 ×
IronLLM-0.6B-Light
GDN hybrid
483
465
428
392
1.23 ×
Qwen3.5-0.8B
GDN hybrid
422
409
380
352
1.20 ×
LFM2-700M
Conv hybrid
508
498
457
416
1.22 ×
Qwen3-0.6B
Full attention
476
402
251
166
2.87 ×
Table 6: Single-user decoding speed over long generations. Decode speed (tokens/s) at position L of a 32K-token generation at batch size 1, i.e. the inverse of the per-token latency (TPOT) averaged over the 512 tokens preceding L , and the growth of TPOT from 1K to 32K. RTX 4090, vLLM 0.17.1, BF16, CUDA graphs, greedy decoding; median of 3 prompts. No speculative decoding.
Benchmark (Metric)
IronLLM-0.6B
IronLLM-0.6B-Light
Qwen3-0.6B 7 7 7 The Qwen3, Qwen3.5, LFM2, and MiniCPM5 results areobtained in non-thinking mode.
LFM2-700M 7
Qwen3.5-0.8B 7
MiniCPM5-1B 7
A.T.
S.E.
A.T.
S.E.
A.T.
S.E.
A.T.
S.E.
A.T.
S.E.
A.T.
S.E.
Long Context
RULER
47.0
703.71
49.3
652.56
125.1
53.57
50.4
463.96
44.3
694.94
117.0
161.06
LongBench v2
173.9
58.62
201.8
52.91
868.1
5.44
1371.6
1.69
652.7
15.01
1064.1
4.27
General Knowledge
MMLU-Pro
345.1
46.01
401.5
35.01
36.9
109.68
675.7
13.26
4652.2
2.68
1731.3
5.93
Table 7: Inference-efficiency comparison between IronLLM-0.6B and representative models. A.T. and S.E. denote average output token length and Score Efficiency, respectively.
Category
Benchmark
General SFT
Domain Expert
TIES Merging
MOPD
Mathematics
GSM8K
61.0
83.9
68.9
78.6
AIME 2026 (Avg@16)
0.6
14.0
0.6
11.0
Code
HumanEval
37.8
70.1
13.4
46.3
MBPP
25.0
57.8
43.8
45.0
LiveCodeBench v6 (Pass@3)
10.9
24.0
14.9
16.6
Instruction Following
IFEval
45.8
77.5
38.5
75.6
Table 8: Capability integration results for IronLLM-0.6B. Domain Expert denotes the independently trained specialist for each domain, while TIES Merging and MOPD produce unified multi-domain models. Bold and underlined values indicate the best and second-best results, respectively.
Figure 8: The conventional MTP architecture. Each prediction depth uses an independent MTP block.
Figure 9: The shared-KV MTP architecture. All prediction depths reuse one MTP block, while keys and values are constructed only at the first depth and shared by subsequent depths.
Architecture
Main Model
MTP Module
Total Parameters
Standard MTP ( K=3 )
653.70M
20.45M × 3
715.05M
Shared-KV MTP ( K=3 )
653.70M
20.45M × 1
674.15M
Table 9: Parameter comparison between three-depth standard MTP and shared-KV MTP. Component counts are approximate.
Benchmark
N
Avg. out len
τ
Decode TPS
Speedup
w/o MTP
K=3
K=1
K=2
K=3
IFEval
81
237
2.48
460
477
0.83 ×
1.02 ×
1.04 ×
GSM8K
99
139
3.42
460
662
0.92 ×
1.27 ×
1.44 ×
MBPP
99
93
3.27
460
629
0.91 ×
1.23 ×
1.37 ×
MATH-500
93
317
3.55
460
678
0.93 ×
1.29 ×
1.48 ×
ShareGPT
87
268
2.82
460
540
0.87 ×
1.11 ×
1.17 ×
Table 10: Decoding speedup of shared-KV X-MTP. Batch size 1, greedy decoding with exact verification by the backbone, so drafts never change the output distribution. N counts the prompts (of 100 sampled per benchmark) whose response ends with EOS in every run, τ is the mean number of tokens committed per backbone step at K=3 , and TPS is the decode throughput with prefill excluded. RTX 4090, vLLM 0.17.1.
Figure 10: The lightweight verification head. At every MTP depth, a shared two-layer MLP consumes the aligned input hidden state and the output hidden state of that depth and emits a scalar acceptance confidence, which gates whether drafting proceeds to the next depth.
Benchmark
Main Model
w/ Ver. Head
Δ
Avg. Accept. Len.
Match Rate
MMLU-Redux
60.35
60.43
+ 0.08
2.41
0.983
IFEval
75.60
72.83
− 2.77
2.23
0.968
GSM8K
78.62
66.26
− 12.36
2.12
0.967
Table 11: Evaluation of the jointly trained verification head with the inference threshold uniformly set to 0.9. Main Model denotes non-speculative decoding with the backbone; Avg. Accept. Len. is the average number of tokens committed per decoding step, and Match Rate is the fraction of confidence-accepted draft tokens that agree with the tokens committed by the backbone, both measured under confidence-gated MTP decoding.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Test Set
IronLLM-0.6B
Orig.
Non-Contam.
Δ
MMLU-Pro
42.1
43.1
+1.0
MMLU-Redux
60.4
61.0
+0.6
C-Eval
47.3
47.0
−0.3
CMMLU
49.4
49.7
+0.3
IFEval
75.6
75.7
+0.1
Appendix
Table 12: Contamination Analysis. Contaminated samples are identified as the union of N-gram overlap detection against the pretraining corpus and embedding-based semantic retrieval against the SFT data.
Figure 11: Pre-training loss curve across the three training stages. The gray trace shows the raw per-step loss and the colored curves show the loss smoothed with a rolling average over 1K steps. The loss decreases smoothly throughout training.
Figure 12: Comparison of CF and MCF evaluation trajectories for two pretraining runs conducted from scratch.
Figure 13: MOPD training curves over the first 10K steps. Left: total optimized loss (dark) and the distillation term LKD alone (light); the warm-up steps, where the total loss peaks at 6.4 , fall outside the plotted y -range. Right: batch-mean verifiable reward. Faint traces show raw per-step values, bold curves an exponential moving average (decay 0.97 ).
Figure 14: GSM8K accuracy of MTP inference w/ verification head (uniform threshold τ=0.9 ) during MOPD training. The accuracy rises steeply in the early training phase and saturates into a stable plateau by the end of the stage.
Figure 15: Effect of the uniform confidence threshold τ on IFEval: accuracy (left), average acceptance length (middle), and confidence–verification match rate (right). Ten uniform thresholds from 0.50 to 0.95 are evaluated.