cs.DCSep 28, 2026

Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines

Authors: Murtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya

Organizations: Quantum Cloud Computing and Distributed Systems (qCLOUDS) Lab, School of Computing and Information Systems, The University of Melbourne, Australia

Abstract

Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding membership removes what they depend on: a swarm holds each part of the model on several peers and routes around a slow one. A bounded session must use every device it admits. Its pipeline advances at the pace of whichever device received a share it cannot serve quickly, so the division has to be right before serving begins. We propose Kafila, whose protocol assembles a ring from behind NATs, preferring direct paths and relaying where traversal fails, while its planner measures each device's memory bandwidth, capacity and reachability, divides the model exactly for a fixed ring order, and places the head, which holds the embedding and output projection, together with that division rather than beforehand. On machines with different capabilities across three fleets, from a shared LAN to five devices spanning two continents, Kafila shortens the slowest pipeline stage by up to 5.2×5.2\times against the even split of pipeline parallelism, as in GPipe, and up to 3×3\times against the memory-proportional split of personal-device inference, as in exo, keeps 75 to 87 per cent of the committed hardware doing work where those divisions fall below half, and serves a model no uniform split can place on the fleet at all. What that is worth to a user depends on how much of a token is computation rather than network. Where the members share a network the same division returns 1.56×1.56\times the throughput of a uniform split and 1.25×1.25\times of a memory-proportional one, and under four concurrent users that lead compounds to 3.2×3.2\times rather than fading, each user served at almost the rate of one.

Figures & tables

Explore similar work

CardsList
  1. Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters

    Apr 7, 2025Zonghang Li, Tao Li, Wenjiao Feng +8LLM Inference OptimizationOffloading

  2. Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models

    Jul 14, 2026Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi +9Large Language Model TrainingPretraining

  3. Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure

    Sep 30, 2026Matias Parij, Pawan Paudel, Tate Berenbaum +1Artificial Intelligence InfrastructureLoad Balancing