Despite its efficiency, there has been little research on the practical aspects required for real-world deployment of on-device AI models, such as the device's CPU utilization and thermal conditions. In this paper, through extensive experiments, we investigate two key issues that must be addressed to deploy on-device models in real-world services: (i) the selection of on-device models and the resource consumption of each model, and (ii) the capability and potential of on-device models for domain adaptation. To this end, we focus on a task of translating live-stream chat messages and manually construct LiveChatBench, a benchmark consisting of 1,000 Korean-English parallel sentence pairs. Experiments on five mobile devices provide a systematic empirical assessment of widely adopted on-device models, highlighting the importance of model selection and deployment constraints when adapting them to specialized tasks and serving a large and heterogeneous user base. We expect that our findings will offer practical insights into both the capabilities and limitations of on-device models for real-world AI service deployment.
Figures & tables
Source (ko)
Target (en)
Background Knowledge
메접 빌드업 지린다
This build-up to quitting MapleStory is insane.
메접 : ’메접’은 온라인 게임 메이플스토리에서 접는다 (그만둔다)는 뜻으로 쓰이는 줄임말입니다.
빌드업 : ’빌드 업(build up)’은 ’쌓아 올리다’, ’점점 증가시키다’라는 뜻으로, 어떤 것을 단계적으로 만들어가는 과정을 의미합니다.
애가 억타 인데
The kid is being forced to play StarCraft, though.
억타 : ’억지로 스타크래프트를 한다’라는 의미. ’억마크’처럼 ’억+(게임 이름)’으로 응용된다.
어제 비챤 님 롤대회가 인기많던데..
Yesterday, VIichan’s League of Legends tournament seemed really popular.
비챤 : 인터넷 방송 플랫폼 "SOOP (숲)"의 스트리머 이름.
Table 1: Examples from the LiveChatBench dataset.
BM25
LLM
BM25 + LLM
Micro Recall ( ↑ )
0.4834
0.7031
0.8107
Table 2: Micro-averaged recall results. The methods compared are BM25, LLM-based entity extraction (using GPT-5.1), and a hybrid BM25+LLM approach.
CPU
GPU
RAM
iPhone 11 Pro
2-core Apple Lightning 2.67 GHz, 4-core Apple Thunder 1.73 GHz
4-core Apple 3rd-generation design GPU architecture, 0,000 MHz
4 GB LPDDR4X SDRAM
iPhone 14 Pro Max
2-core Apple Everest 3.46 GHz, 4-core Apple Sawtooth 2.02 GHz
5-core Apple G14, 1,398 MHz
6 GB LPDDR5 SDRAM
iPhone 16 Pro Max
Dual-core Apple Everest (3rd generation) 4.04 GHz, Quad-core Apple Sawtooth (3rd generation) 2.40 GHz
6-core Apple G17, 1,470 MHz
8 GB LPDDR5X SDRAM
Galaxy S24+
1 × ARM Cortex - X4 3.21 GHz, 2 × ARM Cortex - A720 2.90 GHz, 3 × ARM Cortex - A720 2.60 GHz, 4 × ARM Cortex - A520 at 1.95 GHz
Table 3: Specifications of the mobile devices used in our experiments, including three iOS devices and two Android devices. On iOS devices, we used Apple’s Metal framework for GPU-accelerated inference.
Model Weights
KV Cache
Compute Buffer
Total
Gemma-3-270M
511 MB
192 MB
18 MB
721 MB
Qwen3-0.6B
1137 MB
112 MB
112 MB
1361 MB
Gemma-3-1B
1907 MB
26 MB
192 MB
2125 MB
Table 4: Component-wise memory breakdown. We report the total memory footprint of the Gemma-3-270M, QWEN3-0.6B, and Gemma-3-1B models, decomposed into model weights, KV cache, and compute buffers.
Figure 1: Mobile device temperature results on LiveChatBench across different on-device models.
Figure 2: CPU utilization results on LiveChatBench across different on-device models.
Parameter
Value
n_ctx
512
n_threads
2
n_batch
256
n_uBatch
192
flash_attn_type
On
n_predict
128
Table 5: Model and completion parameters. We employ a unified configuration for all on-device models.
Mobile Device
TTFT (ms)
Runtime (s)
Gemma-3-270M-IT †
iPhone 11 Pro
140.9460
0.4161
iPhone 14 Pro Max
121.4954
0.3412
iPhone 14 Pro Max ‡
74.7435
0.1582
iPhone 16 Pro Max
79.3681
0.2608
iPhone 16 Pro Max ‡
63.3153
0.1598
Table 6: Evaluation of three on-device models on LiveChatBench across five mobile devices. † denotes models trained on the dataset constructed in Section 2 . ‡ indicates results measured in an environment where the mobile device’s GPU was used.
Method
LiveChatBench
FLORES-200
WMT24++
BLEU ( ↑ )
ChrF++ ( ↑ )
FSP ( ↑ )
BLEU ( ↑ )
ChrF++ ( ↑ )
FSP ( ↑ )
BLEU ( ↑ )
ChrF++ ( ↑ )
FSP ( ↑ )
MLKit
0.0489
17.4392
28.2600
0.2556
44.9525
52.5267
0.1737
39.5039
50.3096
Google Translate API
0.1678
34.2974
59.7250
0.4449
60.5056
94.1294
0.3333
53.5668
91.7806
GPT-5.1
0.2679
45.2609
70.2210
0.4113
59.0201
96.8607
0.2998
50.8218
96.2745
Gemma-3-270M-IT
0.0097
5.3872
15.1300
0.0740
20.9606
23.8636
0.0553
17.4603
23.1293
Qwen3-0.6B
0.0017
8.6502
5.9980
0.0261
21.9348
4.8715
0.0221
18.2811
5.7615
Table 7: Comparison of model performance across three different translation datasets. Higher values indicate better translation quality. † denotes models trained on the dataset constructed in Section 2 .
Figure 3: Heatmap of six types of translation errors. Darker colors indicate a higher number of errors.
Figure 4: Comparison of the translation outputs of GPT-5.1 and trained Gemma-3-270M.
Figure 5: Heatmap of six types of translation errors (FLORES-200). Darker colors indicate a higher number of errors.
Figure 6: Heatmap of six types of translation errors (WMT24++). Darker colors indicate a higher number of errors.
Parameter
Value
Total Rank ( r )
64
Scaling Factor ( α )
32
Target Modules
{q, k, v, o, gate, up, down}_proj
Optimizer
AdamW
Warmup Ratio
0.03
Gradient Accumulated Batch
4
Table 8: LoRA training hyperparameters. We employ a unified configuration for all on-device models.
Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales for on-device deployment remain largely unexplored. To close this gap, we present MobileMoE, a family of on-device MoE language models with sub-billion active parameters (0.3-0.9B active and 1.3-5.3B total) that establish a new Pareto frontier for on-device LLMs. We first formulate an on-device MoE scaling law that jointly optimizes MoE architecture under mobile memory and compute constraints, identifying an on-device sweet spot - moderate sparsity with fine-grained and shared experts - that is simultaneously memory and compute-optimal. Building on the derived architectures, we train MobileMoE with a four-stage recipe covering pre-training, mid-training, instruction fine-tuning, and quantization-aware training, all on open-source datasets. Across 14 benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs with 2-4× fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters. To bridge the last mile to mobile deployment, we provide the first efficient MoE inference on commodity smartphones with comprehensive on-device profiling. At comparable INT4 weight memory, MobileMoE-S delivers 1.8-3.8× faster prefill and 2.2-3.4× faster decode than the dense baseline MobileLLM-Pro.
We present ELMOD - Efficient Language Model for On-Device Deployment - a compact (2.7B) German language model designed for efficient inference on resource-constrained hardware. ELMOD was trained on a limited computational budget (55k H100 GPU hours) using exclusively publicly available data. We developed a suite of German-specific data pre-processing, which differ from English-oriented counterparts in their handling of morphological variation, compounding, and orthographic conventions. Furthermore, we introduced a quality filtering and rephrasing step, which increased the instructional quality of the data, improved performance during the annealing phase, and reduced overall compute requirements. Thanks to our architectural model and data choices, including prefiltering, our educational-quality filtering and rephrasal to raise the educational-quality, ELMOD is the strongest performer in its size class (<3B), matching the performance of 7B-parameter models in German.
This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-one inference, low latency, and privacy constraints. These conditions limit the value of optimizations designed for long-context or high-throughput language-model serving. Starting from LMT-60-0.6B, preliminary profiling suggests that vocabulary projection becomes a more important decode-time cost after GGUF quantization reduces the relative cost of Transformer blocks. We replace the original 151k-token vocabulary with a 64k-token subtitle-domain tokenizer, migrate the embedding space, and adapt the model through embedding calibration followed by full supervised fine-tuning. On a fixed 500-example subset of the OpenSubtitles2024 test set, the LocalSubs achieves a 59.2% tie-excluded win rate against Google Translate under GPT-4o pairwise judging. Performance is strongest on short cues and declines as cue length increases. Preliminary Apple M2 Metal measurements on a 64k-vocabulary model show a 1.63× speedup over a 151k-vocabulary profiling baseline. The raw benchmark configuration is incomplete, so the latency result is treated as preliminary.
Tsz-To Wong
National Yang Ming Chiao Tung University, Hsinchu, Taiwan