Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their high resource demands, LLMs are mostly deployed in the cloud. Layer-wise edge-cloud inference lets resource-constrained edge devices contribute computation to LLMs they cannot host in full. However, heterogeneous split points introduce two coupled inefficiencies. First, edge execution and communication create idle gaps between cloud invocations. Second, requests arriving at different model depths cannot be conventionally batched. We present DySCo, a collaborative runtime that keeps KV caches local and introduces dyForward, a model-aware layer-range executor that runs configurable contiguous layer ranges from resident model shards without reloading weights. For multi-edge serving settings, we introduce depth-synchronized batching (DSB), which advances heterogeneous requests to the deepest cut and batches their common suffix. Experiments across heterogeneous devices, two model families, and local and wide-area links show that idle gaps increase the latency of subsequent GPU forward calls even when waiting time is excluded, adding up to 25 ms of additional cloud-side suffix latency per decoding step in our measurements. At an average concurrency of eight, DSB improves throughput by 275% over FIFO, 48% over exact-match batching, and 79% over round-robin interleaving while reducing mean per-session latency. Together, these results show that requests with different edge-cloud splits can reuse resident cloud weights and share batched suffix computation. The artifact repository for this work is publicly available at: https://github.com/Large-scale-Sustainable-Computing-LSC/dysco-artifact
Figures & tables
Fig. 1 : Pipeline and tensor parallelism for one decoding step. Pipeline sharding transfers the cut-layer hidden state, whereas tensor parallelism requires collective communication within every layer.
Fig. 2 : Motivating results from separate single-edge and multi-edge experiments. (a) Controlled idle gaps add cloud suffix latency beyond the injected wait. (b) DSB throughput is less sensitive to cut diversity than exact-match batching.
Fig. 3 : Edge–cloud split-inference architecture with per-device KV caches and a shared cloud request queue.
Fig. 4 : Session initialization, prefill, and autoregressive decode in the edge–cloud inference pipeline.
Fig. 5 : Dynamic layer-range execution across the resident edge and cloud shards.
Phase
Edge → Cloud
Cloud → Edge
Setup
Cut layer and generation configuration
Session acknowledgment
Prefill
Hidden states, attention mask, and position IDs
Token ID
Decode
Hidden state
Token ID
TABLE I : Messages transmitted during normal execution.
Fig. 6 : Depth-synchronized batching. Sessions with different cut layers advance over private ranges to d=maxici and then share one batched forward call over [d,L) . For equal cut layers, DSB reduces to exact-match batching.
Device
GPU mem.
CPU
RAM
Jetson Orin
Shared
Cortex-A78AE
8 GB
RTX A2
16 GB
AMD EPYC 7402P
128 GB
RTX A5000
24 GB
AMD EPYC 7402P
128 GB
RTX A6000
48 GB
2 × AMD EPYC 7402
128 GB
TABLE II: Devices used in the single-edge evaluation.
Fig. 7 : Mean latency breakdowns for Qwen3.5-9B with an A2 edge and an A6000 cloud.
Fig. 8 : Short-input decode latency across edge devices with the cloud fixed to the same A6000.
Fig. 9 : Additional A6000 decode latency after controlled idle gaps. Lines show the median over 31 prompts; shaded regions span the 25th–75th percentiles.
Nˉ
FIFO
Exact-match
Interleave
DSB
throughput (tok/s)
2
58.3
122.6 (+110%)
119.0 (+104%)
129.8 (+122%)
4
58.8
142.6 (+143%)
121.4 (+107%)
186.3 (+217%)
8
58.5
148.3 (+153%)
122.5 (+109%)
219.8 (+275%)
mean latency (s)
2
26.56
7.94
7.20
4.93
TABLE III : Throughput and mean per-session latency under Poisson arrivals at mean concurrency Nˉ (Llama-3.2-3B, WildGPT-4, 128 generated tokens; 16 workers at cut layers 2–9; 15 trials of 24 sessions). Percentages are relative to FIFO; bold denotes a significant best result ( ∣t∣>2.5 ).
Fig. 10 : DSB and exact-match throughput versus shared-tail depth (top) and cut-layer diversity (bottom) for eight Llama-3.2-3B sessions with WildGPT-4 prompts. Each sweep fixes the other variable; error bars show one standard deviation over six trials.
Llama-3.2-3B (28 layers)
Qwen2.5-3B (36 layers)
Tail
Gain
Tail
Gain
20 (71%)
+78.1%
26 (72%)
+66.0%
17 (61%)
+71.6%
22 (61%)
+64.9%
14 (50%)
+63.5%
18 (50%)
+55.4%
11 (39%)
+53.0%
14 (39%)
+51.9%
8 (29%)
+32.7%
10 (28%)
+20.7%
TABLE IV : DSB throughput gain over exact-match batching for eight consecutive cut layers at matched relative tail depths on both models (WildGPT-4, 128 generated tokens; six trials). Parentheses give tail depth as a fraction of model depth; bold denotes ∣t∣>2.5 .
Fleet
Tail
Gain
Homogeneous
8@2
26
−1.3%
8@8
20
+0.1%
8@14
14
−0.7%
One outlier
7@10,1@2
18
+6.3%
TABLE V : DSB gain over exact-match batching by fleet composition ( N=8 , Llama-3.2-3B; six trials). Here, a@c denotes a sessions at cut layer c , Tail denotes L−maxici , and bold indicates ∣t∣>2.5 .
Context
Exact-match
DSB
Gain
64
86.0 ± 0.5
158.1 ± 4.4
+83.9%
128
85.8 ± 0.5
154.1 ± 5.1
+79.7%
256
84.7 ± 0.5
153.2 ± 3.8
+80.7%
512
84.4 ± 0.5
150.5 ± 5.4
+78.2%
1024
83.0 ± 0.5
137.4 ± 4.4
+65.5%
2048
79.4 ± 0.6
112.7 ± 2.4
+42.0%
TABLE VI: Decode throughput versus LongReason context length (Llama-3.2-3B, N=8 , cut layers 1–8, 20-layer shared tail; six trials). Values are mean ± standard deviation; gain is relative to exact-match batching.
N
FIFO
Exact-match
Interleave
DSB
2
61.5
57.2 ( −7% )
102.8 (+67%)
52.7 ( −14% )
4
60.8
97.9 (+61%)
121.6 (+100%)
102.9 (+69%)
6
58.8
107.6 (+83%)
122.7 (+109%)
139.3 (+137%)
8
57.0
112.4 (+97%)
121.8 (+114%)
163.5 (+187%)
TABLE VII: Throughput under synchronized bursts with N sessions at distinct cut layers selected from 9 down to 2 (Llama-3.2-3B; 30 trials with reversed scheduler order). Percentages are relative to FIFO; bold denotes a significant best result ( ∣t∣>2.5 ).
On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference is unfeasible for most consumer and embedded edge devices. This paper presents a privacy-centric edge-cloud collaborative LLM inference framework built on endpoint-authenticated KV cache. Local endpoints handle input preprocessing, embedding computation, adaptive feature optimization, KV cache authentication, speculative decoding and low-dimensional model head calculation, while the cloud conducts authenticated decoder inference, KV cache management, token verification and high-dimensional vocabulary projection. Endpoints fuse partial outputs, apply language-adaptive masking and sample target tokens. All transmitted data and truncated logits are quantized and AES-GCM encrypted for privacy, with core lightweight modules, draft parameters and cache access policies kept local to avoid leakage. The framework supports heterogeneous devices including CPU-only, GPU-equipped and embedded devices via optimized streaming, batching and quantized ONNX deployment. Evaluations demonstrate that the framework reduces per-token latency by up to 46.1% and downlink payloads by up to 67.4% over baseline split inference, retaining comparable performance to full cloud inference.
Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment must address cost efficiency, low latency, and optimal resource utilization. Conventional approaches typically assume that an entire model can be hosted on a single device, which does not hold in many real-world scenarios, particularly in Edge and Fog environments where device resources are constrained. In this paper, we introduce E2LLM, a framework designed to enable efficient LLM deployment in such resource limited settings. Rather than simply partitioning a single model across all available devices, E2LLM replicates the full model across multiple groups of devices (replicas) and applies model parallelism within each replica. Each replica is assigned a specialized role PREFILL or DECODER based on its efficiency in handling input and output tokens. This separation leverages the inherent differences between these two phases of LLM inference. To effectively organize devices, we utilize a Genetic Algorithm to form clusters that maximize system performance. Within each cluster, we apply Dynamic Programming to determine an optimal partitioning strategy that minimizes bottlenecks in model-parallel execution. Experimental results demonstrate that our approach adapts robustly to varying workloads, including scenarios with significant variation in input and output token lengths. Compared to the Splitwise baseline, E2LLM reduces average waiting time by over 50% under high-demand conditions
Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La +3
Department of Informatics University of Oslo Oslo, Norway · Department of Computer Science UiT The Arctic University of Norway Tromsø, Norway
Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with degraded accuracy. Deploying smaller cloud-based big LLMs preserves performance, but at the cost of expensive per-token computation. We present a distributed inference framework, \our{}, that integrates speculative decoding (SD) across edge and cloud. A compact draft model deployed on the edge generates candidate tokens rapidly, and a large verifier model on the cloud validates these tokens in parallel. Accepted tokens are retained, while only rejections trigger verifier correction, substantially reducing the number of cloud queries. Our plug-and-play design shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement. Our approach demonstrates a practical path toward scalable, cost-efficient, and accurate deployment of LLMs in real-world environments. Experimental results across multiple Natural Language Processing tasks using SpecBench and CNN/Dailymail datasets demonstrate that \our{} reduces the cloud model calls by 76% with zero loss in accuracy as compared to the full model.