cs.DCOct 6, 2026

DySCo: Dynamic Sharding for Collaborative Edge-Cloud LLM Inference with Depth-Synchronized Batching

Authors: Jingpo Xu, Paul Joe Maliakel, Ivona Brandic, Shashikant Ilager

Organizations: University of Amsterdam Amsterdam, Netherlands · TU Wien Vienna, Austria

Abstract

Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their high resource demands, LLMs are mostly deployed in the cloud. Layer-wise edge-cloud inference lets resource-constrained edge devices contribute computation to LLMs they cannot host in full. However, heterogeneous split points introduce two coupled inefficiencies. First, edge execution and communication create idle gaps between cloud invocations. Second, requests arriving at different model depths cannot be conventionally batched. We present DySCo, a collaborative runtime that keeps KV caches local and introduces dyForward, a model-aware layer-range executor that runs configurable contiguous layer ranges from resident model shards without reloading weights. For multi-edge serving settings, we introduce depth-synchronized batching (DSB), which advances heterogeneous requests to the deepest cut and batches their common suffix. Experiments across heterogeneous devices, two model families, and local and wide-area links show that idle gaps increase the latency of subsequent GPU forward calls even when waiting time is excluded, adding up to 25 ms of additional cloud-side suffix latency per decoding step in our measurements. At an average concurrency of eight, DSB improves throughput by 275% over FIFO, 48% over exact-match batching, and 79% over round-robin interleaving while reducing mean per-session latency. Together, these results show that requests with different edge-cloud splits can reuse resident cloud weights and share batched suffix computation. The artifact repository for this work is publicly available at: https://github.com/Large-scale-Sustainable-Computing-LSC/dysco-artifact

Figures & tables

Explore similar work

CardsList
  1. E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments

    Jun 2, 2026Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La +3LLM Inference OptimizationEdge Devices

  2. SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

    Aug 13, 2026Divya Jyoti Bajpai, Kishan Kumar Upadhyay, Manjesh Kumar HanawalLLM Inference OptimizationSpeculative Decoding