Federated Learning Aggregation
Momentum
26 papers in the last four weeks, up 117% on the four weeks before. 0.3% of all new papers.
Latest papers 167
Federated Averaging (FedAvg) often degrades under non-IID client data, but it remains unclear whether this degradation reflects the loss of client-learned representations or a failure to use representations that are still present. We study this question mechanistically in sparse client-trained vision models, using dense-model controls to test whether the observed effects depend on sparsity. Our analysis combines class-specific circuit discovery, linear probing of frozen representations, head-only finetuning, and sparse feature dictionaries. Across CNN and ResNet models on CIFAR-10 and Fashion-MNIST, severe label skew can drive some per-class accuracies near zero even when class-specific internal structure remains recoverable. Linear probes substantially outperform the aggregated classifier, head-only finetuning partially restores accuracy, and USAE transfer reveals a largely shared feature basis between IID and non-IID models. Together, these diagnostics suggest that, in our setting, non-IID FedAvg degradation is not fully explained by representational erasure; it also reflects misalignment between preserved internal structure and the final prediction pathway.
FedVideoMAE: Efficient Federated Video Moderation with Differential Privacy and Secure Aggregation
Short-form video moderation is increasingly pushed toward edge and privacy-sensitive settings, where users may intend videos for a limited audience, such as friends or private groups, but sending raw clips to a central server can broaden exposure, consume bandwidth, and add moderation latency. Federated learning can keep videos on device, but unprotected model updates may still leak information, and full-video backbones are expensive to communicate. We present FedVideoMAE, a privacy-preserving federated framework for violence detection that adapts a frozen VideoMAE backbone with lightweight LoRA and prompt parameters. Each training round combines self-supervised masked video reconstruction with client-side differential privacy and pairwise masked aggregation (SA) of adapter updates. Violence labels are held out from federation and used only for downstream evaluation, separating private representation learning from supervised assessment. On RWF-2000, exchanging 5,518,848 trainable parameters instead of the 156,371,328-parameter instantiated pretraining state gives a 28.3x model-state payload ratio. FedVideoMAE reaches 77.25% test accuracy without DP or SA, while accuracy under DP+SA remains in the 65.25-66.00% range. Transfer experiments on RLVS and binary UCF-Crime show similar behavior. These results characterize the privacy-utility trade-off for edge video moderation. Code is available at: https://github.com/zyt-599/FedVideoMAE
Over-the-Air Federated Learning: Rethinking Edge AI Through Signal Processing
Over-the-Air Federated Learning (AirFL) is an emerging paradigm that tightly integrates wireless signal processing and distributed machine learning to enable scalable AI at the network edge. By exploiting wireless superposition over a shared multiple-access channel, AirFL turns simultaneous transmissions of local model updates into an analog aggregate at the receiver, thereby reducing communication latency, bandwidth usage, and energy consumption in wireless aggregation domains. This article develops a design-oriented tutorial view of analog AirFL. We organize existing schemes according to the signal-processing mechanism used to enable AirFL aggregation: transmitter-side channel compensation and power control, receiver-side equalization and high-dimensional processing, or learning-aware aggregation weighting. This viewpoint leads to three representative classes -- CSIT-aware, blind, and weighted AirFL -- and clarifies their assumptions, performance tradeoffs, complexity, and deployment limitations. We further discuss synchronization, digital and hybrid analog-digital realization, and open research directions for integrating AirFL into practical wireless edge-AI systems.
A Nesterov-Accelerated Byzantine-Robust Federated Learning
We investigate robust federated learning, where a group of workers collaboratively train a shared model under the orchestration of a central server in the presence of Byzantine adversaries capable of arbitrary and potentially malicious behaviors. To simultaneously enhance communication efficiency and resilience against such adversaries, we propose a Byzantine-resilient Nesterov-accelerated federated learning (Byrd-NAFL) algorithm. Byrd-NAFL seamlessly integrates Nesterov's momentum into the federated learning process alongside Byzantine-resilient aggregation rules to achieve fast and safe convergence against gradient corruption. We establish a finite-time convergence guarantee for Byrd-NAFL under non-convex and smooth loss functions with relaxed assumptions on the aggregated gradients. Extensive numerical experiments validate the effectiveness of Byrd-NAFL and demonstrate the superiority over existing benchmarks in terms of convergence speed, accuracy, and resilience to diverse malicious attacks.
FedIA: Importance-Aware Aggregation for Domain-Robust Federated Graph Learning
Federated graph learning (FGL) is a natural paradigm for social-media user graphs, where language communities, regional markets, and service boundaries can prevent raw graph pooling. We use the Twitch Gamers networks as the primary live-streaming social-media benchmark, and study a question that is often hidden by representation-level evaluation: after local message passing, what update signals are actually exposed to server aggregation? Through update-space measurements, we identify an aggregation-level failure in which graph-domain clients gradually place salient update signals on less shared parameter coordinates, while message-passing backbones show weaker cross-domain directional compatibility than an MLP control. This update-support fragmentation means that standard averaging can dilute locally important coordinates even when no raw graph data are exchanged. Across five backbones, homophily remains relevant but is not the dominant correlate of support retention; feature, label, and degree discrepancies show stronger associations. These findings indicate that graph-domain shifts damage not only local representations, but also the coordinate support on which aggregation operates. Motivated by this diagnosis, we propose FedIA, a plug-and-play server-side importance-aware aggregation method. Importance Masking selects a shared high-magnitude coordinate support, and Contribution-Aware Momentum Weighting smooths client contributions within that support. FedIA requires no raw graph sharing, no graph-statistics upload, and no auxiliary communication payload, while adding only persistent server state for model coordinates and clients.
FedDAF: Federated Domain Adaptation Using Model Functional Distance
Federated Domain Adaptation (FDA) improves model performance at a target client by collaborating with source clients while preserving data privacy. FDA faces two key challenges: domain shift between source and target data, and limited labeled data at the target, a common constraint when a new site joins a federation before it has accumulated its own labeled data, as in clinical deployments. Most existing methods address domain shift alone, assuming ample target data; those that also tackle data scarcity still fail to prioritize source information according to the target's specific objective. We propose FedDAF, which addresses both challenges through similarity-based aggregation of the global source and target models, using their model functional distance, computed from the angle between their mean gradient fields on target data and normalized via a Gompertz function. The global source model itself is formed using a distance-based weighted average, giving greater weight to source models closer to the target model. Experiments on real-world datasets show FedDAF outperforms existing federated learning (FL), personalized FL, and FDA methods in test accuracy.
FedUHD: Unsupervised Federated Learning using In-Memory Hyperdimensional Computing
Unsupervised federated learning (UFL) enables privacy-preserving distributed training without data labeling, yet practical deployment remains challenging due to non-IID data, high computational and communication costs at edge devices, and sensitivity to communication noise. We propose FedUHD, the first UFL framework based on Hyperdimensional Computing (HDC). On the client side, FedUHD employs kNN-based cluster hypervector removal to mitigate non-IID effects by filtering detrimental local outliers. On the server side, cluster-aware HDC aggregation leverages cluster-level statistics to stabilize learning across heterogeneous clients. To further improve efficiency, we design a compute-in-memory (CIM) accelerator based on a novel phase-change memory (PCM) device, integrated with lightweight ASIC digital modules to execute the client-side HDC pipeline within the accelerator. The intrinsic robustness of HDC to low precision and device variations enables efficient mapping onto analog PCM crossbars, exploiting massive parallelism while minimizing data movement. Experimental results show that FedUHD achieves comparable accuracy to state-of-the-art neural network-based UFL methods across all datasets. On HAR and CIFAR10/100, FedUHD delivers an average 2,239x speedup and 1,542x higher energy efficiency on GPU. In addition, FedUHD reduces communication cost by up to 176x on HAR and CIFAR10/100 and demonstrates greater robustness than Orchestra under communication noise. Compared to GPU implementation of FedUHD, the proposed PCM-based accelerator provides an additional 4.07x speedup and three orders of magnitude higher energy efficiency on average. Furthermore, the results demonstrate the benefit of PCM over RRAM as a CIM substrate.
Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI
Background: Federated learning (FL) enables collaborative training of clinical AI models without centralizing patient data, but adoption is limited by privacy concerns, heterogeneous institutional compliance, and resource disparities; standard differential privacy (DP) applies uniform noise to all clients, penalizing well-compliant or under-resourced institutions. Objective: We introduce a compliance-aware FL framework that adapts DP to institutional compliance, letting lower-compliance sites participate without uniformly penalizing others. Methods: A compliance scoring tool aligned with HIPAA, GDPR, NIST, ISO, and HL7/FHIR maps each client score to a per-step Gaussian noise scale for server-side DP-SGD on a small aggregator dataset. The formal bound applies to the aggregator dataset under a semi-honest aggregator; client-level DP needs secure aggregation (future work). We evaluate five FL strategies on PneumoniaMNIST and BreastMNIST (16 clients, 50 rounds, five seeds); the cumulative aggregator-dataset is about 1434 (Breast) and 513 (Pneumonia) at . Results: Including 12 lower-compliance clients (Experiment 1) versus a compliant-only baseline (Experiment 4) changed BreastMNIST accuracy by +4.5 (FedAvg), +6.8 (FedMedian), +5.2 (FedProx), +1.6 (FedYogi), and -4.1 (FedAdam) percentage points (pooled +2.8 pp; not significant at n=5; up to +17 pp per configuration); compliance-weighted allocation matched uniform server-side DP at equal mean noise (+0.1 pp), carrying no utility penalty, and first-round noise cost 1.3 pp (Breast) and 2.5 pp (Pneumonia, FedAvg). Conclusions: Compliance-weighted server-side DP lets lower-compliance institutions join FL without degrading performance, giving auditable per-site noise control at no utility cost; formal guarantees apply to the aggregator dataset, with client-level DP requiring secure aggregation.
Federated Independent Component Analysis via Spectral Alignment and Robust Aggregation
This paper studies robust estimation for Independent Component Analysis (ICA) in the federated learning setting, where data are locally distributed across clients and may exhibit substantial heterogeneity. The goal is to recover a common global mixing matrix by aggregating local estimators computed by individual clients. The main difficulty is that local ICA estimators are identifiable only up to sign flips and column permutations and may have highly heterogeneous estimation quality. We propose a novel three-step aggregation method that first aligns the permutations across all local estimators via a particular spectral clustering approach, then aligns the signs within each estimated cluster via another tailored spectral approach, and finally applies the geometric median for robust aggregation. The proposed estimator is shown to remain accurate even when a substantial fraction of local estimators are of low quality or inconsistent, as long as each cluster contains a majority of accurate estimators. This contrasts with its analogue based on simple averaging, whose performance is determined by the worst local estimator. In the homogeneous setting, the robust estimator is also shown to be minimax optimal, up to a logarithmic factor, both when the number of clients remains fixed and when it diverges. The theoretical findings are corroborated by simulation studies and a real-data analysis.
Asynchronous Federated Continual Segmentation with Evolving Clients and Label Spaces
Federated learning seeks to foster collaboration among distributed clients while preserving the privacy of their local data. Traditional federated learning methods typically assume a fixed setting, where participating clients, client data, and learning objectives remain unchanged. However, in real-world scenarios, a federation may evolve over time, with changes in both its client composition and target label space. In this evolving federated setting, conventional round-wise model aggregation becomes inflexible, as each federation update requires repeated communication, repeated local computation, and synchronized participation from all accumulated clients. To address this limitation, we propose CA-MMDS, a continual multiple-model distillation framework for federated continual segmentation with asynchronous clients and evolving label spaces. Instead of repeatedly aggregating model parameters from all clients, CA-MMDS maintains a server-side archive of client models and updates the global model through proxy-based distillation from multiple archived local models. When new clients join or existing clients evolve, only the newly added or updated local models need to be uploaded, while unchanged clients can remain offline and continue to contribute through their archived models. This design substantially reduces communication and computation costs while enabling flexible asynchronous cooperation among evolving clients. Using multi-class 3D abdominal CT segmentation as an application task, we demonstrate that CA-MMDS efficiently incorporates evolving client knowledge while achieving competitive segmentation performance.
Approaching the Harm of Gradient Attacks While Only Flipping Labels
Machine learning systems deployed in distributed or federated environments are highly susceptible to adversarial manipulations, particularly availability attacks -- rendering the trained model unavailable. Prior research in distributed ML has demonstrated such adversarial effects through the injection of gradients or data poisoning. In this work, we ask whether comparable degradation is still possible under a substantially more constrained action space: the adversary may only flip a limited number of labels of existing training examples, without modifying features, injecting samples, or directly controlling gradients. We analyze the extent of damage caused by constrained label flipping attacks against distributed learning under mean aggregation -- the dominant baseline in research and production. Focusing on classification problems, (1) we propose a novel formalization of label flipping attacks as a per-round constrained optimization problem, derive a greedy label-selection rule for logistic regression, and empirically evaluate it beyond its derivation setting, including on MLPs and robust aggregators. The rule is provably per-epoch optimal for the attacker under the mean aggregator. (2) Empirically, we show that optimized label flipping can cause substantial accuracy degradation while outperforming random label flipping under similar budgets. (3) We shed light on an interesting interplay between what the attacker gains from more write-access versus what they gain from more flipping budget. (4) Finally, although the attack is derived for mean aggregation, we find that it can transfer empirically to the coordinate-wise median and trimmed mean aggregators, where its effectiveness approaches that of the Little-is-Enough gradient attack. This demonstrates that even highly constrained label-flipping adversaries can pose a significant availability threat to distributed learning.
Influence-Oriented Personalized Federated Learning
Federated learning (FL) is a machine learning paradigm where clients with different behaviors and preferences can learn collaboratively without compromising data privacy. Typical FL methods often rely on fixed weighting for parameter aggregation, thereby neglecting the mutual influence among clients. In practice, clients with similar preferences or backgrounds may provide more useful knowledge to each other, which can be leveraged to improve local performance. However, how to quantify such cross-client influence and how to exploit it for personalized aggregation remain underexplored. To address this gap, we propose an influence-oriented Federated learning framework which quantitatively measures Client-level and Class-level Influence to realize adaptive parameter aggregation for each client (FedC^2I for short). Our core idea is to explicitly model the inter-client influence within an FL system via the well-crafted influence vector and influence matrix. Specifically, FedC^2I incorporate influence vectors to quantify client-level influence, enables clients to selectively acquire knowledge from others, and guides the aggregation of feature representation layers. Meanwhile, the influence matrix captures class-level influence in a more fine-grained manner to achieve personalized classifier aggregation. We evaluate the performance of FedC^2I against existing federated learning methods under non-IID settings, and the results demonstrate the superiority of our method in terms of effectiveness, robustness, and interpretability.
Byzantine-Robust Aggregation for Securing Decentralized Federated Learning
Federated Learning (FL) emerges as a distributed machine learning approach that addresses privacy concerns by training AI models locally on devices. Decentralized Federated Learning (DFL) extends the FL paradigm by eliminating the central server, thereby enhancing scalability and robustness through the avoidance of a single point of failure. However, DFL faces significant challenges in optimizing security, as most Byzantine-robust algorithms proposed in the literature are designed for centralized scenarios. In this paper, we present a novel Byzantine-robust aggregation algorithm to enhance the security of Decentralized Federated Learning environments, coined WFAgg. This proposal handles adverse conditions and strengthens the robustness of dynamic decentralized topologies at the same time by employing multiple filters to identify and mitigate Byzantine attacks. Experimental results demonstrate the effectiveness of the proposed algorithm in maintaining model accuracy and convergence in the presence of various Byzantine attack scenarios, outperforming state-of-the-art centralized Byzantine-robust aggregation schemes (such as Multi-Krum or Clustering). These algorithms are evaluated on an IID image classification problem in both centralized and decentralized scenarios.
FedReview: Review and Dispose Poisoned Updates without Validation Datasets or Historic Knowledge
Federated learning has emerged as a decentralized approach for training high-performance models without accessing user data. Despite its effectiveness, it is vulnerable to poisoning attacks, where malicious users manipulate the global model by uploading poisoned updates. In this paper, we propose FedReview, a review-based mechanism to identify and dispose the potential poisoned updates in federated learning. Under FedReview, the server randomly assigns a subset of clients as reviewers to evaluate model updates on their training datasets in each round. The reviewers rank the updates based on evaluation results and estimate the number of low-quality updates as potential poisoned ones. Based on the review reports, the server applies a majority voting mechanism to aggregate rankings, which tolerates wrong rankings from malicious reviewers and guides the removal of suspicious updates during model aggregation. In contrast to prior works such as FLTrust, FedReview does not require a server-side validation dataset or prior knowledge of clients, allowing flexible client participation. Extensive experiments demonstrate that FedReview enables the server to learn a well-performing global model in adversarial environments.
Reinforcement Federated Learning Method Based on Adaptive OPTICS Clustering
Federated learning is a distributed machine learning technology, which realizes the balance between data privacy protection and data sharing computing. To protect data privacy, feder-ated learning learns shared models by locally executing distributed training on participating devices and aggregating local models into global models. There is a problem in federated learning, that is, the negative impact caused by the non-independent and identical distribu-tion of data across different user terminals. In order to alleviate this problem, this paper pro-poses a strengthened federation aggregation method based on adaptive OPTICS clustering. Specifically, this method perceives the clustering environment as a Markov decision process, and models the adjustment process of parameter search direction, so as to find the best clus-tering parameters to achieve the best federated aggregation method. The core contribution of this paper is to propose an adaptive OPTICS clustering algorithm for federated learning. The algorithm combines OPTICS clustering and adaptive learning technology, and can effective-ly deal with the problem of non-independent and identically distributed data across different user terminals. By perceiving the clustering environment as a Markov decision process, the goal is to find the best parameters of the OPTICS cluster without artificial assistance, so as to obtain the best federated aggregation method and achieve better performance. The reliability and practicability of this method have been verified on the experimental data, and its effec-tiveness and superiority have been proved.
Class-wise Contribution Estimation via Logit Maximization for Federated Learning
Federated learning (FL) enables collaborative learning of computer vision models, where privacy and regulatory constraints prevent centralizing data across devices or organizations. However, practical FL deployments often exhibit severe class imbalance and label skew, causing standard aggregation protocols to overfit dominant clients and degrade minority-class performance. We propose a data-free, class-wise contribution estimation and aggregation framework based on logit maximization (CELM) that does not require sharing raw data, client metadata, or auxiliary public datasets. The FL server probes client updates to obtain class-wise evidence scores and assembles a cross-client evidence matrix, which quantifies both per-class competence and class coverage. Using this matrix, we compute contribution weights that upweight clients providing strong, discriminative evidence for underrepresented classes. The resulting aggregation is stable due to simplex constraints and momentum smoothing, and it remains compatible with standard FL training pipelines. We evaluate the approach on representative vision benchmarks under controlled non-IID and pathological label splits, demonstrating that CELM-based aggregation improves robustness to imbalance and statistical heterogeneity, while yielding better performance without requiring any additional data exchange.
Byzantine-Robust Federated Fire Detection with a Rotating Coordinator
We study the application of federated learning (FL) to indoor fire detection. Such fire-detection systems use edge cameras that record sensitive footage which cannot easily be collected at a central server. Existing federated solutions leave three practical obstacles unaddressed: limited uplink bandwidth, Byzantine (malicious or faulty) clients, and unconditional trust in a single, permanently fixed aggregation server. Our main contributions address all three. In particular, we provide (i) a curated indoor fire-detection dataset assembled from eight public sources; (ii) an edge-deployable detector whose model updates are compressed up to 10 times with only a small loss in balanced accuracy; and (iii) a semi-decentralized Byzantine-robust FL method that combines history-aware aggregation with a rotating coordinator, evicting stealthy attacks that per-round filters miss while removing the fixed-server single point of failure. On the held-out test set the rotating-coordinator method matches its fixed-server counterpart in accuracy and detection speed, and a physically distributed six-node cloud deployment confirms feasibility.