Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure
Organizations: Not Community Labs Inc. · Intel Corporation
Abstract
We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers. An operator-run certificate authority handles admission and fleet management outside the inference path. Three serving modes share one interface: whole-model execution on one node, load-balanced replicas, and pipeline-sharded chains using the compilation and speculative decoding mechanism of our companion paper. Optional KV-cache mobility reuses compatible conversation prefixes after a routing move, with cold recomputation on a miss. Signed response receipts and hash-chained logs support provenance and audit. A three-node Phi-3.5-mini NPU testbed delivered 3.10x the response throughput of its one-node configuration under ten concurrent requests; a separate four-node deployment recorded 4.06x the throughput of direct single-node serving. Paired latency observations, runtime measurements, and internal functional checks characterize the tested configurations. We compare Cascadia with IBM, Nutanix, VMware, and HPE platforms on deployment footprint, hardware requirements, scheduling, scaling, licensing, and trust, using vendor documentation. The paper repository provides benchmark scripts, curated measurements, and a claim-to-evidence map.
Figures & tables
| Metric | Direct OVMS | Cascadia 4-node | Result |
|---|---|---|---|
| TTFT p50, single request | 1438 ms | 1307 ms | observed medians |
| Aggregate throughput (10 conc.) | 0.104 req/s | 0.422 req/s | |
| Wall time, 100 requests | 16 min | 4 min | |
| p50 latency under load | 87.2 s | 12.7 s | lower |
| Reported success (600+ requests) | 100% | 100% | internal benchmark |
| Nodes | Concurrency | Responses/s | p50 under load | Relative rate |
|---|---|---|---|---|
| 1 | 10 | 0.222 | 45.0 s | |
| 2 | 10 | 0.484 | 18.9 s | |
| 3 | 10 | 0.690 | 12.2 s |
| Quantity | Target | Measured | Source |
|---|---|---|---|
| Revocation propagation, mesh-wide | 60 s | 18–20 s | internal admission trials, Apr 2026 |
| Fourth node reaches mesh-ready (peers known) | n/a | 14 s | internal mesh-join observation |
| Receipt presence (stream / non-stream) | n/a | 100% / 100% | e7 , 2026-06-08 |
| Receipt length (base64); /v1/models p50 | n/a | 272 chars; 137 ms | e7 , 2026-06-08 |
| Workload (Lunar Lake, 32 GB UMA) | iGPU-only | ILP placement | Note |
|---|---|---|---|
| Yi-1.5-9B fp16 (16.45/16.48 GiB iGPU) | 2.03 tok/s | 2.91 tok/s | ; embedding to CPU |
| Qwen2.5-32B INT4 (comfortable fit) | 1.97 tok/s | 1.27 tok/s | iGPU-only wins; ILP overfits |
| SOLAR-10.7B fp16 (over-nominal) | 1.31 tok/s | 0.79 tok/s | transparent UMA spill unmodeled |
| K2.6 expert dispatch | tok/s | vs. |
|---|---|---|
| Top- (router-faithful) | 0.105 | baseline |
| Top- | 0.325 |
| Dimension | IBM (watsonx / OpenShift AI) | Nutanix (Enterprise AI) | VMware (Private AI Fdn.) | HPE (Private Cloud AI) | Cascadia |
|---|---|---|---|---|---|
| Minimum footprint | OpenShift cluster (3 control-plane HA) + Software Hub + GPU workers; KServe; Serverless / Service Mesh in serverless mode [ 57 , 58 , 59 ] | 3 control-plane + 3 worker K8s nodes [ 44 ] ; Cisco CVD entry: 4 HCI nodes, 2 L40S each [ 14 ] | Full VCF stack (SDDC Mgr., vCenter, NSX, Supervisor); 3 GPU-enabled hosts [ 9 ] | Turnkey rack; 3 dedicated control nodes + GreenLake cloud mgmt. plane [ 22 ] | AI PCs; no dedicated tier; CA off-path |
| Accelerator floor | NVIDIA A100 / H100 / L40S class; RHEL AI min. L4 24 GB; Gaudi 3 tech preview [ 55 ] | NVIDIA-only list (L40S, A100, H100, H200) [ 46 ] ; CPU mode: Xeon AMX, 10B [ 42 ] | vGPU-capable datacenter NVIDIA 3 hosts; MIG unsupported with NIM [ 9 ] | G1: 4 L40S small production config; 2 H100 NVL developer system [ 22 ] | CPU + iGPU + NPU already in each AI PC |
| Control plane | Kubernetes mandatory; operator stack; license metering in-cluster | Kubernetes mandatory (CNCF); Prism for HCI layer | VCF private-cloud control plane mandatory [ 9 ] | 3 control nodes/rack; mgmt. plane in HPE cloud [ 22 ] | No dedicated request scheduler; CA handles fleet management |
| Scaling model | Replica serving via KServe [ 65 ] ; distributed execution depends on runtime configuration | Replica serving via NIM/KServe; NIM also supports distributed execution [ 49 ] | vLLM/NIM serving; deployment and networking are managed by VCF [ 7 ] | Managed NIM serving on rack-scale GPU configurations [ 22 ] | Replicas and load-balanced chains; client-device pipeline mechanism [ 5 ] ; disk-streamed MoE |
| Licensing meter | Per-VPC on-prem; 67 VPC listed $643,200/yr [ 4 ] ; per-accelerator add-on [ 54 , 61 ] ; per-token/GPU-hr SaaS [ 27 ] | Per GB of aggregate GPU vRAM [ 45 , 12 ] atop per-core platform licenses | Per-core VCF subscription + per-GPU NVIDIA AIE [ 9 , 8 , 11 , 50 ] | GreenLake subscription (3/5-yr terms) + per-GPU NVAIE [ 22 , 50 ] | Closed-source mesh; Apache-2.0 runtime; existing client hardware |
| Air gap / edge | Disconnected OpenShift (heavyweight); Fusion HCI racks [ 23 ] | Dark-site bundles; NKP Metal still K8s on servers [ 47 ] | VCF 9.1 fleet ops incl. air-gapped / sovereign [ 10 ] | G2 air-gapped configuration [ 22 ] | Static binaries + local models; disconnected is the default shape |