Deploying LLMs on edge devices presents serious technical challenges. Memory elasticity is crucial for edge devices with unified memory, where memory is shared and fluctuates dynamically. Existing solutions suffer from either poor transition granularity or high storage costs. We propose FlexQuant, a novel elasticity framework that generates an ensemble of quantized models, providing an elastic hosting solution with 31x more deployment options, 15x granularity improvement, and 10x storage reduction compared to SoTA methods. FlexQuant works with most quantization methods and creates a family of trade-off options under various storage limits through our pruning method. It brings great performance and flexibility to the edge deployment of LLMs.
Figures & tables
Fig. 1: Perplexity comparison of different elastic hosting methods of quantized Llama 27 B model. It compares FlexQuant with the baseline method when they are using two quantization methods, ExLlamaV2 and AnyPrecision.
Fig. 2: An example of FlexQuant’s tree search for EQM ensemble. The shown search process has two exploitation stems and three exploration branches.
Fig. 3: Perplexity comparison between Base-Ex, Base-AP, FQ-Ex, and PFQ-Ex at different pruning rate. Results for AnyPrecision on Llama 3 8B is omitted due to lack of support in their implementation. The footprint range is different due to differences in parameter count and model architecture.
Fig. 4: Perplexity vs pruning rate at varying memory footprint bounds for FQ-Ex on Llama 1, Llama 2 and Llama 3
Fig. 5: Downstream perplexity comparison of quantized Llama models between Base-Ex FQ-Ex and PFQ-Ex at different pruning rate.
Task
arcChallenge
arcEasy
hellaswag
piqa
winogrande
Policy
Base-Ex
FQ-Ex
PFQ-Ex
Base-Ex
FQ-Ex
PFQ-Ex
Base-Ex
FQ-Ex
PFQ-Ex
Base-Ex
FQ-Ex
PFQ-Ex
Base-Ex
FQ-Ex
PFQ-Ex
Storage
177 GB
17 GB
13 GB
177 GB
17 GB
13 GB
177 GB
17 GB
13 GB
177 GB
17 GB
13 GB
177 GB
17 GB
13 GB
Memory Footprint (GB)
3.0
37.40
-1.40
-1.40
68.98
+1.73
+1.73
51.47
-0.60
-0.80
75.63
+0.15
+0.54
66.61
+0.24
-1.10
3.5
39.42
+2.04
+1.08
72.05
+2.57
+1.77
53.21
+3.02
+1.87
76.22
+1.14
+0.60
66.69
+2.22
+1.34
4.0
41.38
+1.19
+1.11
75.01
+0.35
+0.20
55.81
+0.76
+0.55
77.04
+0.76
+0.71
68.27
+1.10
+0.40
4.5
43.17
-0.26
-0.43
75.46
+0.00
-0.25
55.81
+0.93
+0.77
77.75
+0.26
+0.00
69.37
+0.32
-0.39
TABLE I: Downstream task accuracy comparison between Base-Ex, FQ-Ex and PFQ-Ex with 50% pruning.
Fig. 6: Hyperparameters’ impact on the overall search algorithm’s performance. In the left and middle plots, we sweep between different numbers of stems and the number of branches. The right plot shows the impact of the middle model. In the middle and right plots, the lighter the color is, the better the quality of the search result.
Fig. 7: First token latency and decoding throughput overhead of FlexQuant framework. The Baseline is Base-Ex, indicated by the grey bar. The FlexQuant is FQ-Ex, indicated by the light blue bar.
Fig. 8: Memory usage collected on a Samsung S23 Ultra, using the Android 14 operating system. Three collections of workload traces are collected based on the number of active apps on the cellphone. Their memory usage histograms are presented from left to right.
Fig. 9: Model quality comparison under different workload scenarios. ”AP with 6.0GB” or ”FQ with 6.0GB” means AnyPrecision or FlexQuant deployed on a device with 6GB of total system memory. The lighter the color is, the better the model quality.
Fig. 10: LLM launch speed comparison between FlexQuant (FP-Ex) and AnyPrecision (Base-AP). The larger the speedup is, the better the performance.
Deploying large language models (LLMs) is challenging due to their significant memory and computational requirements. While some methods address this by developing small or tiny language models from scratch, these approaches demand extensive GPU training. Compressing pre-trained LLMs for edge devices offers a compelling alternative. Beyond pruning and quantization, Neural Architecture Search (NAS) enables effective compression, yet prior NAS approaches often limit the search space and decouple architecture from quantization. We introduce a differentiable NAS framework that explores the entire space and jointly optimizes architectural configurations alongside mixed-precision quantization for linear layers of LLMs. Experiments demonstrate superior accuracy-latency trade-offs: our models achieve up to 1.4x faster inference than sequential NAS-then-quantization baselines at comparable accuracy, or up to 6% higher average accuracy across seven reasoning tasks at equivalent latency.
Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1
UiT The Arctic University of Norway · University of Oslo, Norway
Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment must address cost efficiency, low latency, and optimal resource utilization. Conventional approaches typically assume that an entire model can be hosted on a single device, which does not hold in many real-world scenarios, particularly in Edge and Fog environments where device resources are constrained. In this paper, we introduce E2LLM, a framework designed to enable efficient LLM deployment in such resource limited settings. Rather than simply partitioning a single model across all available devices, E2LLM replicates the full model across multiple groups of devices (replicas) and applies model parallelism within each replica. Each replica is assigned a specialized role PREFILL or DECODER based on its efficiency in handling input and output tokens. This separation leverages the inherent differences between these two phases of LLM inference. To effectively organize devices, we utilize a Genetic Algorithm to form clusters that maximize system performance. Within each cluster, we apply Dynamic Programming to determine an optimal partitioning strategy that minimizes bottlenecks in model-parallel execution. Experimental results demonstrate that our approach adapts robustly to varying workloads, including scenarios with significant variation in input and output token lengths. Compared to the Splitwise baseline, E2LLM reduces average waiting time by over 50% under high-demand conditions
Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La +3
Department of Informatics University of Oslo Oslo, Norway · Department of Computer Science UiT The Arctic University of Norway Tromsø, Norway
Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of Natural Language Processing (NLP) tasks, but their high computational and memory demands pose significant challenges for deployment on resource-constrained edge devices. Existing approaches to model compression and optimization often rely on coarse-grained pruning or quantization, which can compromise accuracy or require re-training and fine-tuning. In this work, we introduce SelectInfer, a neuron-level optimization framework that enables efficient LLM inference on edge devices through selective neuron loading and computation. By profiling and identifying both task-specific and general-purpose neurons using an offline LLM profiler, SelectInfer implements two key optimizations: selective loading, which reduces memory footprint by selectively loading a subset of neurons that were identified to be most important during the offline stage, and selective computation, which dynamically computes only the most relevant neurons at runtime. Evaluation across multiple datasets shows that SelectInfer achieves significant reductions in memory footprint and computation while preserving task performance, making it a practical step towards enabling LLM deployment on edge devices
Huzaifa Shaaban Kabakibo, Eric Schniedermeyer, Artem Burchanow +1
Department of Computer Networks, Paderborn University, Paderborn, Germany