Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines
Organizations: Quantum Cloud Computing and Distributed Systems (qCLOUDS) Lab, School of Computing and Information Systems, The University of Melbourne, Australia
Abstract
Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding membership removes what they depend on: a swarm holds each part of the model on several peers and routes around a slow one. A bounded session must use every device it admits. Its pipeline advances at the pace of whichever device received a share it cannot serve quickly, so the division has to be right before serving begins. We propose Kafila, whose protocol assembles a ring from behind NATs, preferring direct paths and relaying where traversal fails, while its planner measures each device's memory bandwidth, capacity and reachability, divides the model exactly for a fixed ring order, and places the head, which holds the embedding and output projection, together with that division rather than beforehand. On machines with different capabilities across three fleets, from a shared LAN to five devices spanning two continents, Kafila shortens the slowest pipeline stage by up to against the even split of pipeline parallelism, as in GPipe, and up to against the memory-proportional split of personal-device inference, as in exo, keeps 75 to 87 per cent of the committed hardware doing work where those divisions fall below half, and serves a model no uniform split can place on the fleet at all. What that is worth to a user depends on how much of a token is computation rather than network. Where the members share a network the same division returns the throughput of a uniform split and of a memory-proportional one, and under four concurrent users that lead compounds to rather than fading, each user served at almost the rate of one.
Figures & tables
| System | Workload | A slow or absent participant | Connectivity graph | Per-stage cost obtained from | Ordering by reachability | Head placed with the division |
|---|---|---|---|---|---|---|
| Petals [ 6 ] | LLM inference | Routed around at request time | Complete | Throughput measured at runtime | ||
| SWARM [ 7 ] / Parallax [ 8 ] | LLM inference | Rescheduled at runtime | Complete, unequal bandwidth | Throughput measured at runtime | ||
| GPipe [ 11 ] | Training | Does not arise; deployment is provisioned | Complete, uniform speed | Layer count | ||
| exo [ 12 ] | LLM inference | Slows the ring; not addressed | Complete, on one network | Device memory | ||
| PipeEdge [ 13 ] / PipePar [ 14 ] | Encoder inference; training | Does not arise; deployment is provisioned | Complete, unequal speed | A profile taken with the weights resident | ||
| Kafila (this work) | LLM inference | Cannot be avoided, so it is planned for | Incomplete and unequal | A bandwidth probe, before any weight is fetched |
| Fleet | Device | Mem | Bw (GB/s) | Site |
| LAN | L40S | 24 | 512–529 | Melbourne, AU |
| (4 dev) | L40 | 24 | 367–453 | Melbourne, AU |
| A40 | 12 | 377–439 | Melbourne, AU | |
| M3 Pro | 18 | 89–124 | Melbourne, AU | |
| US Central | A100 | 80 | 915–959 | Des Moines, US |
| (5 dev) | RTX 5090 | 32 | 849–1047 | Chicago, US |
| Initiator | Responder | Rule predicts | Traversal achieves |
|---|---|---|---|
| permissive | permissive | direct | direct |
| permissive | ordinary | direct | direct |
| permissive | restrictive | direct | direct |
| ordinary | permissive | direct | direct |
| ordinary | ordinary | direct | direct |
| ordinary | restrictive | relay | direct |
| Model | Allocation | Imbal. | Bneck | Util. | tok/s |
|---|---|---|---|---|---|
| max/min | ms | ||||
| LAN , 12% network overhead | |||||
| Qwen3-8b | Kafila | 1.46 | 11.3 | 86% | 21.5 0.2 |
| Mem-prop. (exo) | 4.01 | 26.0 | 48% | 17.2 3.8 | |
| Uniform (GPipe) | 5.14 | 36.2 | 41% | 13.7 1.3 | |
| US Central , 38 to 62% network overhead | |||||