Despite its efficiency, there has been little research on the practical aspects required for real-world deployment of on-device AI models, such as the device's CPU utilization and thermal conditions. In this paper, through extensive experiments, we investigate two key issues that must be addressed to deploy on-device models in real-world services: (i) the selection of on-device models and the resource consumption of each model, and (ii) the capability and potential of on-device models for domain adaptation. To this end, we focus on a task of translating live-stream chat messages and manually construct LiveChatBench, a benchmark consisting of 1,000 Korean-English parallel sentence pairs. Experiments on five mobile devices provide a systematic empirical assessment of widely adopted on-device models, highlighting the importance of model selection and deployment constraints when adapting them to specialized tasks and serving a large and heterogeneous user base. We expect that our findings will offer practical insights into both the capabilities and limitations of on-device models for real-world AI service deployment.
Figures & tables
Source (ko)
Target (en)
Background Knowledge
메접 빌드업 지린다
This build-up to quitting MapleStory is insane.
메접 : ’메접’은 온라인 게임 메이플스토리에서 접는다 (그만둔다)는 뜻으로 쓰이는 줄임말입니다.
빌드업 : ’빌드 업(build up)’은 ’쌓아 올리다’, ’점점 증가시키다’라는 뜻으로, 어떤 것을 단계적으로 만들어가는 과정을 의미합니다.
애가 억타 인데
The kid is being forced to play StarCraft, though.
억타 : ’억지로 스타크래프트를 한다’라는 의미. ’억마크’처럼 ’억+(게임 이름)’으로 응용된다.
어제 비챤 님 롤대회가 인기많던데..
Yesterday, VIichan’s League of Legends tournament seemed really popular.
비챤 : 인터넷 방송 플랫폼 "SOOP (숲)"의 스트리머 이름.
Table 1: Examples from the LiveChatBench dataset.
BM25
LLM
BM25 + LLM
Micro Recall ( ↑ )
0.4834
0.7031
0.8107
Table 2: Micro-averaged recall results. The methods compared are BM25, LLM-based entity extraction (using GPT-5.1), and a hybrid BM25+LLM approach.
CPU
GPU
RAM
iPhone 11 Pro
2-core Apple Lightning 2.67 GHz, 4-core Apple Thunder 1.73 GHz
4-core Apple 3rd-generation design GPU architecture, 0,000 MHz
4 GB LPDDR4X SDRAM
iPhone 14 Pro Max
2-core Apple Everest 3.46 GHz, 4-core Apple Sawtooth 2.02 GHz
5-core Apple G14, 1,398 MHz
6 GB LPDDR5 SDRAM
iPhone 16 Pro Max
Dual-core Apple Everest (3rd generation) 4.04 GHz, Quad-core Apple Sawtooth (3rd generation) 2.40 GHz
6-core Apple G17, 1,470 MHz
8 GB LPDDR5X SDRAM
Galaxy S24+
1 × ARM Cortex - X4 3.21 GHz, 2 × ARM Cortex - A720 2.90 GHz, 3 × ARM Cortex - A720 2.60 GHz, 4 × ARM Cortex - A520 at 1.95 GHz
Table 3: Specifications of the mobile devices used in our experiments, including three iOS devices and two Android devices. On iOS devices, we used Apple’s Metal framework for GPU-accelerated inference.
Model Weights
KV Cache
Compute Buffer
Total
Gemma-3-270M
511 MB
192 MB
18 MB
721 MB
Qwen3-0.6B
1137 MB
112 MB
112 MB
1361 MB
Gemma-3-1B
1907 MB
26 MB
192 MB
2125 MB
Table 4: Component-wise memory breakdown. We report the total memory footprint of the Gemma-3-270M, QWEN3-0.6B, and Gemma-3-1B models, decomposed into model weights, KV cache, and compute buffers.
Figure 1: Mobile device temperature results on LiveChatBench across different on-device models.
Figure 2: CPU utilization results on LiveChatBench across different on-device models.
Parameter
Value
n_ctx
512
n_threads
2
n_batch
256
n_uBatch
192
flash_attn_type
On
n_predict
128
Table 5: Model and completion parameters. We employ a unified configuration for all on-device models.
Mobile Device
TTFT (ms)
Runtime (s)
Gemma-3-270M-IT †
iPhone 11 Pro
140.9460
0.4161
iPhone 14 Pro Max
121.4954
0.3412
iPhone 14 Pro Max ‡
74.7435
0.1582
iPhone 16 Pro Max
79.3681
0.2608
iPhone 16 Pro Max ‡
63.3153
0.1598
Table 6: Evaluation of three on-device models on LiveChatBench across five mobile devices. † denotes models trained on the dataset constructed in Section 2 . ‡ indicates results measured in an environment where the mobile device’s GPU was used.
Method
LiveChatBench
FLORES-200
WMT24++
BLEU ( ↑ )
ChrF++ ( ↑ )
FSP ( ↑ )
BLEU ( ↑ )
ChrF++ ( ↑ )
FSP ( ↑ )
BLEU ( ↑ )
ChrF++ ( ↑ )
FSP ( ↑ )
MLKit
0.0489
17.4392
28.2600
0.2556
44.9525
52.5267
0.1737
39.5039
50.3096
Google Translate API
0.1678
34.2974
59.7250
0.4449
60.5056
94.1294
0.3333
53.5668
91.7806
GPT-5.1
0.2679
45.2609
70.2210
0.4113
59.0201
96.8607
0.2998
50.8218
96.2745
Gemma-3-270M-IT
0.0097
5.3872
15.1300
0.0740
20.9606
23.8636
0.0553
17.4603
23.1293
Qwen3-0.6B
0.0017
8.6502
5.9980
0.0261
21.9348
4.8715
0.0221
18.2811
5.7615
Table 7: Comparison of model performance across three different translation datasets. Higher values indicate better translation quality. † denotes models trained on the dataset constructed in Section 2 .
Figure 3: Heatmap of six types of translation errors. Darker colors indicate a higher number of errors.
Figure 4: Comparison of the translation outputs of GPT-5.1 and trained Gemma-3-270M.
Figure 5: Heatmap of six types of translation errors (FLORES-200). Darker colors indicate a higher number of errors.
Figure 6: Heatmap of six types of translation errors (WMT24++). Darker colors indicate a higher number of errors.
Parameter
Value
Total Rank ( r )
64
Scaling Factor ( α )
32
Target Modules
{q, k, v, o, gate, up, down}_proj
Optimizer
AdamW
Warmup Ratio
0.03
Gradient Accumulated Batch
4
Table 8: LoRA training hyperparameters. We employ a unified configuration for all on-device models.