cs.DCAug 12, 2026

User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling

Authors: Alfreds LapkovskisAli BeikmohammadiSindri MagnússonPraveen Kumar Donta

Organizations: Department of Computer Systems and Sciences, Stockholm University SE-106 91 Stockholm, Sweden

Abstract

Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources absorb increasing demand without proportional growth in centralized infrastructure. To capture stochastic and dynamic interactions among users, resources, tasks, and policies, we develop a high-dimensional generative Markov model with structured temporal factorization. The model supports simulation and provides a foundation for task scheduling and QoS-aware resource allocation optimization. We evaluate the system across user populations, resource capacities, and centralized and distributed scheduling policies. Simulations show that distributed scheduling becomes increasingly advantageous as the user population grows, improving request completion and P99 latency while substantially reducing dedicated resource consumption. These results demonstrate the feasibility of user-assisted collaborative inference for infrastructure-efficient autoscaling.

Explore similar work

CardsList
  1. Brief Announcement: Generative Markov Model for Distributed Computing Systems

    Jun 2, 2026Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon +1Distributed OptimizationMarkov