Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their high resource demands, LLMs are mostly deployed in the cloud. Layer-wise edge-cloud inference lets resource-constrained edge devices contribute computation to LLMs they cannot host in full. However, heterogeneous split points introduce two coupled inefficiencies. First, edge execution and communication create idle gaps between cloud invocations. Second, requests arriving at different model depths cannot be conventionally batched. We present DySCo, a collaborative runtime that keeps KV caches local and introduces dyForward, a model-aware layer-range executor that runs configurable contiguous layer ranges from resident model shards without reloading weights. For multi-edge serving settings, we introduce depth-synchronized batching (DSB), which advances heterogeneous requests to the deepest cut and batches their common suffix. Experiments across heterogeneous devices, two model families, and local and wide-area links show that idle gaps increase the latency of subsequent GPU forward calls even when waiting time is excluded, adding up to 25 ms of additional cloud-side suffix latency per decoding step in our measurements. At an average concurrency of eight, DSB improves throughput by 275% over FIFO, 48% over exact-match batching, and 79% over round-robin interleaving while reducing mean per-session latency. Together, these results show that requests with different edge-cloud splits can reuse resident cloud weights and share batched suffix computation. The artifact repository for this work is publicly available at: https://github.com/Large-scale-Sustainable-Computing-LSC/dysco-artifact
Figures & tables
Fig. 1 : Pipeline and tensor parallelism for one decoding step. Pipeline sharding transfers the cut-layer hidden state, whereas tensor parallelism requires collective communication within every layer.
Fig. 2 : Motivating results from separate single-edge and multi-edge experiments. (a) Controlled idle gaps add cloud suffix latency beyond the injected wait. (b) DSB throughput is less sensitive to cut diversity than exact-match batching.
Fig. 3 : Edge–cloud split-inference architecture with per-device KV caches and a shared cloud request queue.
Fig. 4 : Session initialization, prefill, and autoregressive decode in the edge–cloud inference pipeline.
Fig. 5 : Dynamic layer-range execution across the resident edge and cloud shards.
Phase
Edge → Cloud
Cloud → Edge
Setup
Cut layer and generation configuration
Session acknowledgment
Prefill
Hidden states, attention mask, and position IDs
Token ID
Decode
Hidden state
Token ID
TABLE I : Messages transmitted during normal execution.
Fig. 6 : Depth-synchronized batching. Sessions with different cut layers advance over private ranges to d=maxici and then share one batched forward call over [d,L) . For equal cut layers, DSB reduces to exact-match batching.
Device
GPU mem.
CPU
RAM
Jetson Orin
Shared
Cortex-A78AE
8 GB
RTX A2
16 GB
AMD EPYC 7402P
128 GB
RTX A5000
24 GB
AMD EPYC 7402P
128 GB
RTX A6000
48 GB
2 × AMD EPYC 7402
128 GB
TABLE II: Devices used in the single-edge evaluation.
Fig. 7 : Mean latency breakdowns for Qwen3.5-9B with an A2 edge and an A6000 cloud.
Fig. 8 : Short-input decode latency across edge devices with the cloud fixed to the same A6000.
Fig. 9 : Additional A6000 decode latency after controlled idle gaps. Lines show the median over 31 prompts; shaded regions span the 25th–75th percentiles.
Nˉ
FIFO
Exact-match
Interleave
DSB
throughput (tok/s)
2
58.3
122.6 (+110%)
119.0 (+104%)
129.8 (+122%)
4
58.8
142.6 (+143%)
121.4 (+107%)
186.3 (+217%)
8
58.5
148.3 (+153%)
122.5 (+109%)
219.8 (+275%)
mean latency (s)
2
26.56
7.94
7.20
4.93
TABLE III : Throughput and mean per-session latency under Poisson arrivals at mean concurrency Nˉ (Llama-3.2-3B, WildGPT-4, 128 generated tokens; 16 workers at cut layers 2–9; 15 trials of 24 sessions). Percentages are relative to FIFO; bold denotes a significant best result ( ∣t∣>2.5 ).
Fig. 10 : DSB and exact-match throughput versus shared-tail depth (top) and cut-layer diversity (bottom) for eight Llama-3.2-3B sessions with WildGPT-4 prompts. Each sweep fixes the other variable; error bars show one standard deviation over six trials.
Llama-3.2-3B (28 layers)
Qwen2.5-3B (36 layers)
Tail
Gain
Tail
Gain
20 (71%)
+78.1%
26 (72%)
+66.0%
17 (61%)
+71.6%
22 (61%)
+64.9%
14 (50%)
+63.5%
18 (50%)
+55.4%
11 (39%)
+53.0%
14 (39%)
+51.9%
8 (29%)
+32.7%
10 (28%)
+20.7%
TABLE IV : DSB throughput gain over exact-match batching for eight consecutive cut layers at matched relative tail depths on both models (WildGPT-4, 128 generated tokens; six trials). Parentheses give tail depth as a fraction of model depth; bold denotes ∣t∣>2.5 .
Fleet
Tail
Gain
Homogeneous
8@2
26
−1.3%
8@8
20
+0.1%
8@14
14
−0.7%
One outlier
7@10,1@2
18
+6.3%
TABLE V : DSB gain over exact-match batching by fleet composition ( N=8 , Llama-3.2-3B; six trials). Here, a@c denotes a sessions at cut layer c , Tail denotes L−maxici , and bold indicates ∣t∣>2.5 .
Context
Exact-match
DSB
Gain
64
86.0 ± 0.5
158.1 ± 4.4
+83.9%
128
85.8 ± 0.5
154.1 ± 5.1
+79.7%
256
84.7 ± 0.5
153.2 ± 3.8
+80.7%
512
84.4 ± 0.5
150.5 ± 5.4
+78.2%
1024
83.0 ± 0.5
137.4 ± 4.4
+65.5%
2048
79.4 ± 0.6
112.7 ± 2.4
+42.0%
TABLE VI: Decode throughput versus LongReason context length (Llama-3.2-3B, N=8 , cut layers 1–8, 20-layer shared tail; six trials). Values are mean ± standard deviation; gain is relative to exact-match batching.
N
FIFO
Exact-match
Interleave
DSB
2
61.5
57.2 ( −7% )
102.8 (+67%)
52.7 ( −14% )
4
60.8
97.9 (+61%)
121.6 (+100%)
102.9 (+69%)
6
58.8
107.6 (+83%)
122.7 (+109%)
139.3 (+137%)
8
57.0
112.4 (+97%)
121.8 (+114%)
163.5 (+187%)
TABLE VII: Throughput under synchronized bursts with N sessions at distinct cut layers selected from 9 down to 2 (Llama-3.2-3B; 30 trials with reversed scheduler order). Percentages are relative to FIFO; bold denotes a significant best result ( ∣t∣>2.5 ).