May 12, 2026 · cs.ITJ/K move · Enter open · S save
Nan Xue, Shengkang Chen, Zhiyong Chen, Jiangchao Yao+3
Cooperative Medianet Innovation Center, Shanghai Jiao Tong University, Shanghai 200240, China · Department of Broadband Communication, Pengcheng Laboratory, Shenzhen 518000, China · Future Network of Intelligent Institute (FNii), the Chinese University of Hong Kong (Shenzhen), Shenzhen 518172, China · Meta, Menlo Park, CA 94025 USA
As large language models (LLMs) move from centralized clouds to mobile edge environments, efficient serving must balance latency, energy consumption, and accuracy under constrained device-edge resources. Query-level routing between lightweight on-device models and stronger edge models provides a flexible mechanism to navigate this trade-off. However, existing routers are designed for centralized cloud settings and optimize token-level costs, failing to capture the dynamic latency and energy overheads in wireless edge deployments. In this paper, we formulate mobile edge LLM routing as a deployment-constrained, cost-aware decision problem, and propose CR^2, a two-stage device-edge routing framework. CR^2 decouples a lightweight on-device margin gate from an edge-side utility selector for deferred queries. The margin gate operates on frozen query embeddings and a user-specified cost weight to predict whether local execution is utility-optimal relative to the best edge alternative under the target operating point. We further introduce a conformal risk control (CRC) calibration procedure that maps each operating point to an acceptance threshold, enabling explicit control of the marginal false-acceptance risk under the full-information utility reference. Experiments on the routing task show that CR^2 closely matches a full-information reference router using only device-side signals before deferral. Compared with strong query-level baselines, CR^2 consistently improves the deployable accuracy-cost Pareto frontier and reduces normalized deployment cost by up to 16.9% at matched accuracy.