Federated LLM fine-tuning enables large models to be adapted using private and geographically distributed data at the network edge, creating recurring and deadline-sensitive communication workloads across access and transport networks. This challenge is particularly relevant in mobile RANs, where wireless variability, mobility, and device heterogeneity cause model updates to arrive asynchronously. Although these updates belong to the same learning round and share a common destination and deadline, conventional transport networks treat them as independent device-originated flows, hiding their underlying structure and limiting the ability to efficiently provision transport resources. This mismatch is particularly problematic for optical circuit switching and all-photonics transport, which benefit from predictable and schedulable traffic demands. We argue that future RANs should act as learning-aware traffic shapers by exposing the communication structure of distributed model adaptation to the transport layer. Through in-network aggregation at the gNB, asynchronous UE updates can be transformed into fewer aggregate transfers with bounded size and delivery requirements. Once shaped in this way, federated LLM traffic becomes a suitable candidate for selectively provisioned optical connectivity, where high-capacity paths can be established during aggregate-transfer windows and released between learning rounds. The resulting architecture combines the flexibility of packet-based mobile access with dynamically provisioned optical capacity, illustrating a broader approach for coordinating distributed AI workloads across programmable access and transport networks.
Figures & tables
Fig. 1 : (Left) Federated LLM fine-tuning over the RAN: (a) learning workflow and (b) traffic timeline. (Right) Learning-aware gNB aggregation: (c) processing pipeline and (d) transport timeline.
Fig. 2 : Learning-aware APN/OCS scheduling of gNB aggregate transfers.
Fig. 3 : Effect of gNB in-network aggregation on model accuracy and round duration.
Fig. 4 : gNB-to-parameter-server traffic with conventional forwarding and in-network aggregation.
Fig. 5 : Packet-versus-OCS completion time as a function of the number of active gNBs for small, medium, and large PEFT adapter aggregates.
Emilio Paolini received his Ph.D. degree in Emerging Digital Technologies, cum laude, from Scuola Superiore Sant’Anna, Pisa, Italy, in 2024. He is currently an Assistant Professor with the TeCIP Institute, Scuola Superiore Sant’Anna, Pisa, Italy. His research interests include AI-enabled NextG networks, programmable data planes, and federated learning. He received the GTTI Ph.D. Thesis Award for the Best Italian Ph.D. in Communication Technologies in 2024.
Table 6
Andrea Pinto received his BS and MS in computer engineering from the University of Naples Federico II. He graduated with a Ph.D. in Computer Science from Saint Louis University in 2026, specializing in computer networks, distributed systems, and scalable machine learning systems. His research focuses on distributed training and bandwidth-aware orchestration to alleviate network bottlenecks.
Table 7
Flavio Esposito is an Associate Professor and Graduate Coordinator in the Department of Computer Science at Saint Louis University, where he also serves as a Fellow of the Research Institute. He received his Ph.D. in Computer Science from Boston University. His research focuses on edge computing, programmable and virtualized networks, network security, CPS, and applied artificial intelligence.
Table 8
Luca Valcarenghi is a Full Professor at the Scuola Superiore Sant’Anna of Pisa, Italy, since 2024. He received the Ph.D. from UTD in 2001. Dr. Valcarenghi received a Fulbright Research Scholar Fellowship in 2009 and a JSPS ”Invitation Fellowship Program for Research in Japan (Long Term)” in 2013. He coordinated, as a PI or local PI, several National and International projects. His main research interests are optical networks design, analysis, and optimization; energy efficiency in communications networks; optical access networks; zero touch network and service management; 5G technologies and beyond.
Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data. However, in mobile deployments, the training wall-clock is often dominated by straggler-limited uplink communication under heterogeneous bandwidth, intermittent participation, and non-IID client data. Although parameter-efficient fine-tuning (PEFT) methods such as LoRA and QLoRA reduce local memory and trainable parameters, repeated transmission of adapter updates remains a major bottleneck. We propose Fed-FSTQ, a semantic-sensitivity-aware communication-control primitive for communication-efficient federated LLM fine-tuning. Fed-FSTQ uses a lightweight token-level Fisher proxy to estimate semantic sensitivity, couples token-guided sparsification with mixed-precision adapter-update quantization, and allocates higher communication fidelity to semantically load-bearing evidence while suppressing redundant transmission. The method is drop-in compatible with standard federated PEFT pipelines and requires no change to the server aggregation rule. Experiments on multilingual QA and medical QA under non-IID partitions show that Fed-FSTQ reduces cumulative uplink traffic required to reach a fixed quality threshold by 46-fold relative to a Fed-LoRA baseline and improves straggler-limited wall-clock time-to-accuracy by 52%. Under the corrected Controlled LTE-20Mbps accounting, Fed-FSTQ reduces per-round time from 414.60s to 67.29s and reduces per-round energy from 839.20J to 146.28J, yielding a 6.16-fold speedup. On NVIDIA Jetson-class edge devices, Fisher-guided token reduction also yields up to a 1.55-fold inference speedup, demonstrating deployability under tight resource constraints.
Changyu Li, Shuanghong Huang, Jiashen Liu +5
School of Computing and Information Technology, Great Bay University, Dongguan, China · Beijing Institute of Technology, Beijing, China · University of Warwick, Coventry, U.K. +3
Large Language Models (LLMs) have significantly propelled the advancement of edge intelligence and have been widely deployed across various scenarios, including autonomous driving, industrial inspection, and personalized IoT services. However, the collaborative adaptation of LLMs on edge devices continues to face formidable challenges due to strict data privacy constraints, highly heterogeneous computing and communication resources, and the non-independent and identically distributed (non-IID) nature of local data. Federated Fine-Tuning (FFT) enables the collaborative optimization of distributed models without exposing raw data. Yet, traditional synchronous aggregation suffers from a severe straggler effect, resulting in high system latency and low resource utilization. Existing asynchronous federated learning methods are predominantly designed for small-to-medium-scale models and struggle to address the specific challenges inherent in LLM fine-tuning namely, model drift caused by stale updates, aggravated client drift stemming from data heterogeneity, and aggregation fairness imbalance resulting from the dominance of fast clients. To address these issues, this paper proposes AlignFed, an asynchronous federated fine-tuning framework for LLMs tailored to heterogeneous edge environments. AlignFed employs a lightweight multi-stage semantic alignment mechanism comprising three core modules: version-aware update grouping, cross-version semantic alignment based on a mini-batch calibration set, and fairness-aware aggregation that integrates both update freshness and client participation frequency. This framework effectively mitigates cross-version model drift and client drift while enhancing aggregation fairness, thereby achieving stable and efficient asynchronous federated optimization in scenarios characterized by high heterogeneity and significant update staleness.
Yan Wang, Ziyi Gao, Rui Wang
Department of Computer and Communication Engineering, University of Science and Technology Beijing, Beijing 100083, China
Federated Split Learning has been identified as an efficient approach to address the computational resource constraints of clients in classical federated learning, while guaranteeing data privacy for distributed model training across data owners. However, it faces some critical challenges when such a training strategy meets large language models (LLMs) for fine-tuning. Such challenges include setting the cutlayer adaptively across different clients to address the data and device heterogeneity issues, which affect the system performance significantly. In addition, efficiently reducing the communication overhead during the fine-tuning procedure is also another challenge. No work tries to address these challenges. To bridge this gap, we propose SplitTF, an adaptive federated split learning system for LLMs fine-tuning. SplitFT enables different clients to set different cut layers according to their computation resources and trained model performance. SplitFT also proposes to reduce the LoRA rank in cutlayer to reduce the communication overhead. In addition to simulating the heterogeneous data in real-world applications for our proposed split federated learning system, we propose a length-based Dirichlet approach to divide the training data into different clients. Extensive experimental results show that our proposed approach outperforms the state-of-the-art approach for fine-tuning time efficiency and model performance based on various popular benchmarks.
Yimeng Shan, Zhaorui Zhang, Sheng Di +3
Hong Kong Polytechnic University Hong Kong · Argonne National Labratory USA · University of California, Merced USA +1