cs.AISep 29, 2026

IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence

Authors: Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, +9 more

Organizations: Robotics Foundation Model Team, Xpeng Inc.

Abstract

We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference

    Sep 2, 2026Renyuan Liu, Yuyang Leng, Kaiyan Liu +7LLM Inference OptimizationStreaming

  2. SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs

    Jul 20, 2026Huzaifa Shaaban Kabakibo, Eric Schniedermeyer, Artem Burchanow +1LLM Inference OptimizationNeurons