Upholding Robustness in Federated Learning: Trends, Emerging Strategies, and Research Opportunities
Organizations: Department of Computer Science and Engineering, BITS Pilani Dubai Campus, Dubai, UAE · Department of Engineering, University of Palermo, Palermo, Italy · Department of Computer Science, Missouri University of Science and Technology, Rolla, USA
Abstract
While Federated Learning (FL) has been widely adopted for protecting user privacy in machine learning, it remains vulnerable to various robustness challenges, including performance-impairment risks, information-stealing threats, and aggregation vulnerabilities. This work offers a holistic synthesis of FL robustness along three tightly coupled angles: (i) a threat-centric view of robustness that categorizes the multifaceted attack surfaces, (ii) a structured taxonomy of robust aggregation strategies distinguishing outcome-centric approaches from security-centric strategies, and (iii) a layered taxonomy of defensive strategies. We rigorously examine current evaluation practices for FL robustness and identify major applications and open research challenges to guide future research.
Figures & tables
| Role | Involvement in FL |
|---|---|
| Participant | • Corrupt or replace local model updates. • Observe the global model. • Control local training (optimize loss function and hyperparameters (learning rate, local epochs, batch size)). • Coordinate attacks to corrupt or influence the global model. |
| Server | • Directly inspect or manipulate global model parameters (but has no access to train or test data). • Inspect local model updates of participants in the absence of secure aggregation mechanisms. • Inject an attack over the aggregation mechanism. |
| External | • Intercept communication between entities. • Perform inference attacks. • Create a malicious model replica. • End users of the deployed service can act as adversaries (they have access to the final trained model). |
| Class | Adversarial Role | Specific Target | Execution Rounds | AK | |||
|---|---|---|---|---|---|---|---|
| Participant | Server | Data | Model | One | Many | ||
| Targeted | |||||||
| Untargeted | |||||||
| Inference (P) | |||||||
| Inference (M) | |||||||
| Inference (C) | |||||||
| Paper | Threat | Description | Non-IID data | Persistent | Similarity space | Evaluation metrics | Dataset | Remarks |
|---|---|---|---|---|---|---|---|---|
| ( Peri et al., 2020 ) , ( Rong et al., 2022 ) , ( Yang et al., 2023a ) , ( Xia et al., 2023 ) | CLP attacks | Inject poison into samples but retain original labels, ensuring imperceptibility but embedding malicious traits. | Negative Cosine Similarity (NCS) | Accuracy, Attack Success Rate (ASR), Peak Signal-to-Noise Ratio (PSNR) and Structural SIMilarity (SSIM) | MNIST, Fashion MNIST, CIFAR-10, ImageNet | Requires large-scale data to be effective | ||
| ( Xu et al., 2022b ) , ( Tolpegin et al., 2020 ) | DLP attacks | Introduce incorrect labels to data, creating misclassification risks upon model incorporation. | Euclidean Distance | Misclassification Rate, Attack Impact | MNIST, CIFAR-100, KDD, Amazon | More detectable than clean-label poisoning | ||
| ( Sun et al., 2024a ) , ( Alsereidi et al., 2024 ) , ( Jere et al., 2020 ) , ( Zhang et al., 2021a ) , ( Psychogyios et al., 2023 ) | PSG with GANs | Use adversarially trained GANs to create poisoned samples that mimic real data, disrupting global model accuracy. | Frechet Inception Distance, Wasserstein Distance | Model Accuracy, Detection Rate | PlantVillage, CelebA, ImageNet, Tiny-ImageNet | Requires high computational resources | ||
| ( Sastry and Oore, 2020 ) , ( Winkens et al., 2020 ) , ( Ren et al., 2020 ) | EDP attacks | Insert samples from outside the original input distribution to impair model integrity. | Jensen-Shannon Divergence (JSD) | Model Deviation, ASR | OpenImages, MNIST, CIFAR-10 | Limited by dataset shift detection methods | ||
| ( Ranjan et al., 2022a ) , ( Ranjan et al., 2022b ) , ( Tuor et al., 2021 ) , ( Cao et al., 2022 ) , ( Sun et al., 2021a ) , ( Li et al., 2024a ) | Colluding attacks | Multiple controlled clients collude to amplify attack effects, often seen in Sybil attacks. | Cosine Similarity | ASR, Model Deviation | MNIST, CIFAR-10 | Harder to execute with robust aggregation techniques | ||
| ( Hossain et al., 2021 ) , ( Fang et al., 2020 ) | RNI attacks | Inject random noise into gradients or models to disrupt model convergence. | Kullback-Leibler (KL) divergence | Model Convergence, Error Rate | FEMNIST, MNIST, CIFAR-10 | Limited by robust gradient clipping methods |
| Paper | Threat | Description | Non-IID data | Persistence | Similarity space | Evaluation metrics | Dataset | Remarks |
|---|---|---|---|---|---|---|---|---|
| ( Huynh et al., 2024 ) | Single pattern attack | Adversarial clients inject the same malicious pattern into the model. | Model weight similarity | ASR , Accuracy Drop | CIFAR-10, MNIST | Limited adaptability in diverse settings. | ||
| ( Bagdasaryan et al., 2020 ) , ( Abad et al., 2023 ) | Model weight modification | Modify weights to embed hidden patterns influencing specific tasks while preserving accuracy. | Parameter space similarity | ASR, Accuracy | CIFAR-10, ImageNet | Hard to detect as attacked models resemble benign ones. | ||
| ( Xie et al., 2020a ) , ( Gong et al., 2022 ) | Multi-backdoor attack | Multiple coordinated adversaries inject diverse local patterns or segments of a global pattern. | Model update clustering | ASR, False Positive Rate | CIFAR-10, Fashion-MNIST | High complexity makes detection challenging. | ||
| ( Liu et al., 2020 ) | Feature-partioned attack | Backdoor embedded in feature-partitioned FL without label manipulation. | Gradient similarity analysis | ASR, Accuracy | Purchase-100, FEMNIST | Gradient aggregation mitigates attack effectiveness. | ||
| ( Chen et al., 2020 ) | Federated meta-learning attack | Backdoor persists even after meta-training and fine-tuning on clean data. | Task-specific weight similarity | ASR, Model Accuracy | Omniglot, Mini-ImageNet | Reducing fine-tuning effectiveness is a challenge. | ||
| ( Wang et al., 2020b ) | Edge-case backdoor attack | Targeting underrepresented input data to induce misclassification. | Adversarial loss detection | ASR, Robust Accuracy | CIFAR-10, GTSRB | Hard to detect due to its selective nature. |
| Paper | Threat | Description | Non-IID data | Persistence | Similarity space | Evaluation metrics | Dataset | Remarks |
|---|---|---|---|---|---|---|---|---|
| ( Wang et al., 2020a ) , ( Vangala et al., 2022 ) | MITM attack | Disrupt communication between the server and participants, creating a single point of failure. | Communication delay | Packet loss, Communication time | CIFAR-10, MNIST | Exploits communication bottlenecks to intercept updates | ||
| ( Yao et al., 2018 ) | Channel bottleneck exploitation | Exploit server-client communication bottlenecks by reducing bandwidth, introducing delays, and interference. | Latency, bandwidth reduction | Model accuracy, Convergence time | CIFAR-10, MNIST | Can lead to severe performance degradation | ||
| ( Hitaj et al., 2023 ) | Covert channel through advanced encoding | Embed hidden messages in model parameters using encoding techniques, utilizing weights and gradient updates. | Gradient modification, encoding | Model update consistency, Hidden data rate | WikiText-2, CIFAR-10, MNIST, ESC-50 | Covert data transmission in model weights | ||
| ( Costa et al., 2022 ) | Covert channel through input sample poisoning | Create a covert communication channel by encoding data bits into changes observed during federated training rounds. | Data poisoning, input modification | Model accuracy, Detection time | MNIST, CIFAR-10 | Affects model training with undetectable input manipulation | ||
| ( Zhang et al., 2023a ) | DDoS attack | Target server and network traffic to prevent clients from connecting to the server. | Network congestion, client disconnection | Connection uptime, Service disruption | EMNIST, MNIST, CIFAR-10 | Prevents system operation and model updates | ||
| ( Cao et al., 2022 ) , ( Fung et al., 2020 ) | Training inflation (Sybil attacks) | Inflate malicious participants to disrupt the training process. | Identity falsification | Training effectiveness, Model accuracy | MNIST, CIFAR-10, Reddit, HAR, KDDCup | Disrupt FL process with fake clients |
| Paper | Threat | Attack description | Non-IID data | Persistence | Simlarity method | Evaluation metrics | Dataset | Remarks |
|---|---|---|---|---|---|---|---|---|
| ( He et al., 2024 ) | Membership inference (Passive) | Infer whether a data point is a part of the training set by observing model updates. | Cosine similarity | Accuracy of membership prediction | CIFAR-10, MNIST, FMNIST, CIFAR-100 | Can infer membership without active involvement. | ||
| ( Gu et al., 2022 ) | Membership inference (Active) | Tamper the FL model to gain insights into the training data of other participants. | Gradient ascent | Loss reduction rate, Accuracy of inference | CIFAR-10, MNIST | Active attack increases the risk by direct manipulation. | ||
| ( Hu et al., 2024 ) , ( Hu et al., 2021 ) | Source inference | Identify the source participant by analyzing model updates, exposing information about the data owner. | Gradient analysis | Identity recovery accuracy, Success rate of source identification | CIFAR-10, MNIST, CHMNIST | Risk of revealing the participant’s identity. | ||
| ( Suri et al., 2022 ) | Subject inference | Target data owners rather than records, aiming to infer their membership using a trained model. | Loss analysis | Subject inference accuracy | MNIST, Fashion-MNIST | Requires model access post-training for the attack’s effectiveness. | ||
| ( Zhang et al., 2021a ) | Poisoning and inference | Combine poisoning techniques with inference methods to manipulate training data and infer information about participants. | Gradient poisoning | Data manipulation accuracy, Inference accuracy | CIFAR-10, MNIST | More potent when combined with poisoning attacks, allowing for model manipulation. | ||
| ( Arevalo et al., 2024 ) , ( Mothukuri et al., 2021 ) | Attribute inference | Target specific features of participants’ data to infer sensitive information without accessing the full data. | Gradient analysis | Attribute inference accuracy | CIFAR-10, MNIST | Challenging to detect the attacker’s presence. |
| Attacks | Attack Target | Vulnerable Component | Complexity | Type | ||||
| C | S | M | D | Comm | ||||
| Clean-Label Poisoning | Training Data | H | A | |||||
| Dirty-Label Poisoning | Training Data | M | A | |||||
| Poisoned Samples Generation | Training Data | H | A | |||||
| Model Poisoning | Global Model | H | A | |||||
| Backdoor Attacks | Global Model | H | A | |||||
| Category | Paper | Methods | Key Features | Advantages | Drawbacks |
| Client model evaluation | ( Cao et al., 2021 ) , ( Pillutla et al., 2022 ) , ( Park et al., 2021 ) | Cosine similarity based evaluation of gradients, Geometric median aggregation, Entropy-based client weighting. | Prioritizes gradients with higher cosine similarity, Robust aggregation via geometric median, Personalization via entropy-based weighting. | Enhances robustness against malicious clients, Reduces impact of outliers. | Computational overhead in gradient similarity calculations, Requires additional validation metrics, May not generalize to all non-IID settings. |
| Quantum secure aggregation | ( Chehimi and Saad, 2022 ) , ( Javeed et al., 2024 ) , ( Yang et al., 2022 ) , ( Zhang et al., 2022b ) | Quantum bit representation of model parameters, Entangled qubits for model aggregation, Post-quantum secure protocol using homomorphic pseudorandom generator. | Secure aggregation using qubits, Low computational complexity, Post-quantum security, Compatible with different model architectures. | High resilience against attacks, Ensures privacy without sacrificing efficiency. | Requires specialized quantum hardware, Limited scalability in classical systems, Not widely adopted due to early-stage development. |
| Dynamic fusion of local models | ( Lee et al., 2021 ) , ( Wu et al., 2023 ) , ( Sun et al., 2022 ) | Dynamic fusion of local models based on time windows, Adaptive aggregation strategies, Temporal weight decay strategies. | Excludes stale models, Implements deadline-based aggregation. | Handles stale model updates effectively, Improves aggregation efficiency. | Performance depends on accurate time-window selection, High communication overhead for frequent updates, Increased computation for temporal decay management. |
| Statistical features | ( Peng et al., 2024 ) , ( Wang et al., 2024b ) | Trimmed Mean, Median-Krum, Multi-Krum aggregation. | Median-Krum joint method, Robust aggregation under adversarial conditions, Effective accuracy and convergence under malicious conditions. | Robust against adversarial manipulations, Convergence under malicious attacks, Effective under non-IID settings | May discard useful model updates, computationally expensive, May not perform well in extreme data heterogeneity scenarios. |
| Adaptive client selection | ( Du et al., 2023 ) , ( Sultana et al., 2022 ) , ( Wan et al., 2022 ) | Multi-armed bandit strategy, Gradient update norms, Radial-basis functions, Adaptive client selection based on communication capacity. | Multi-armed bandit approach, Dynamic client evaluation, Balancing exploration vs. exploitation, Focus on gradient norms and model alignment. | Dynamic selection based on exploration-exploitation trade-off, Efficient client evaluation to reduce communication overhead. | Selection bias may impact global model accuracy. |
| Training function optimization | ( Li et al., 2019 ) , ( Andrew et al., 2021 ) , ( Zhao et al., 2024 ) | Regularization in loss function, Dynamic clipping value, Huber loss function. | Loss function regularization, Dynamic adaptation for non-IID scenarios, Huber loss extension for optimal robustness. | Adaptive to non-IID scenarios, Improves robustness of convergence. | Requires careful tuning of regularization parameters. |
| Key Challenge | Underlying Causes | Impact on Aggregation | Existing Solutions | Evaluation Metrics | Trade-offs Involved | Gaps in Research |
| Statistical heterogeneity | Non-IID data distribution, Device diversity, Varying client participation, Variability in local datasets. | Divergence in local models affects global model convergence and performance. | Bayesian non-parametric methods, Neuron matching, Probabilistic federated neural matching. | Accuracy, Model convergence rate, Gradient similarity. | Increased computation for better adaptability requires careful tuning. | Generalizing methods for complex neural networks, Adaptive sampling, Meta-learning, Federated transfer learning, Hybrid approaches for improved convergence. |
| Fairness and bias mitigation | Demographic biases in models, Disproportionate impact of non-IID data. | Biased model updates, reducing overall model fairness and generalization. | FairFL, Adaptive sampling, Bias-aware aggregation strategies. | Model fairness metrics (e.g., demographic parity, equal opportunity). | Increased computational costs and require additional fairness constraints in optimization. | Advances in Federated debiasing techniques, Fairness-aware aggregation. |
| Bottlenecks in communication | High number of clients, Limited bandwidth and connectivity constraints, Frequent communication rounds, Latency Constraints, Communication Overhead. | Slows down model updates and convergence | AirComp-based FL (Over-the-air computation), Intelligent reflective surfaces (IRS), Multi-relay techniques, Gradient sparsification, Quantization, Local update compression, 6G integration. | Latency, Communication overhead, Model convergence speed. | Accuracy loss, Higher infrastructure cost. | 6G for higher efficiency and scalability, AI-driven network optimization, Edge caching. |
| Secure aggregation | Model poisoning, Inference attacks, Sybil attacks, Data poisoning, Backdoor attacks. | Risk of poisoning and inference attacks compromise FL security. | Robust aggregation, Anomaly detection TEE, Cryptographic filtering, Blockchain. | Attack resistance, Privacy preservation, Adversarial robustness. | Trade-offs in model utility, Higher computational and storage costs for improved security. | Enhancing blockchain security for decentralized FL, Secure multi-party computation (SMC), Decentralized consensus mechanisms. |
| Robustness and efficiency tradeoff | High computational and communication costs, Resource-constrained edge devices. | Balancing security, accuracy, and resource constraints, Delay in model convergence. | Adaptive aggregation techniques, Sparsification and quantization, Cryptographic solutions. | Accuracy loss, Model convergence time, Computational cost, Computational overhead. | Model robustness and computational overhead, Security, and aggregation speed. | Dynamic aggregation methods that optimize performance under varying conditions, Trade-off-aware federated optimization strategies, lightweight security mechanisms, and robustness under real-time FL conditions. |
| Quantum aggregation | Instability due to quantum errors, Quantum noise, error correction, and limited availability of quantum hardware. | Faster aggregation, Higher security via quantum cryptographic techniques. | Quantum secure multi-party computation, Quantum-enhanced HE, Hybrid quantum-classical FL aggregation models. | Aggregation speed, Computational complexity, Noise resilience in quantum operations, Robustness against quantum attacks. | Quantum and classical processing, Stability and computational speed, Scalability. | Stable quantum aggregation, PQC- methods, Error mitigation techniques, Hybrid quantum-classical models. |
| Category | Paper | Methods | Key features | Evaluation metrics | Datasets | Advantages | Drawbacks |
| Against training-data manipulation | ( Xia et al., 2023 ) , ( Xu et al., 2024a ) , ( Andrew et al., 2021 ) , ( Park et al., 2023 ) , ( Zhu et al., 2023 ) | Advanced preprocessing, local data filtering, enhanced VAE for property division or variation control, DP on data or gradients, global KD to filter malicious client knowledge, adversarial distillation against backdoors. | Reduces poisoning and inference risk, provides noise-based privacy, and filters poisoned samples before aggregation. | ASR, property inference AUC, reconstruction success, and clean accuracy. | MNIST, CIFAR-10/100, EMNIST | Protects data and intermediate representations, reduces attack surface before aggregation, preserves utility via KD/distillation. | DP/KD/VAEs add overhead, utility loss if over-tuned, need auxiliary data, hyperparameter sensitivity. |
| Against malicious model updates | ( Wang et al., 2024b ) ( Wang et al., 2021a ) , ( Cao et al., 2021 ) , ( Peng et al., 2024 ) , ( Wang et al., 2020b ) , ( Yang et al., 2023b ) , ( Zhang et al., 2021c ) , ( Zhang and Li, 2024 ) , ( Chen et al., 2024 ) , ( Kalapaaking et al., 2022 ) , ( Ranathunga et al., 2022 ) , ( Ozfatura et al., 2021 ) , ( Marnissi et al., 2024 ) , ( Wang et al., 2023 ) , ( Jiang and Borcea, 2023 ) , ( Xu et al., 2024a ) , ( Xu et al., 2022a ) ( Xu et al., 2024b ) , ( Sun et al., 2021b ) , ( Ma et al., 2022b ) , ( Zhang and Hu, 2023 ) | Robust aggregation rules, similarity or trust-based filtering, truth-discovery, HE-based, variance-reduced + DP, autoencoder-based anomaly filtering, sparsification, watermarking, blockchain/smart-contracts. | Server or client-side update inspection, down-weight or drop suspicious gradients, and cryptographic traceability for updates. | Accuracy under attack, ASR, detection precision/recall, convergence, watermark detection rate, and blockchain latency. | MNIST, CIFAR-10/100, EMNIST, FEMNIST, synthetic datasets | Broad coverage of model-poisoning and byzantine behaviors, tolerate high malicious ratios, can recover from polluted rounds; reduces single point of failure, integrity, and ownership guarantees. | Bounded assumption on adversarial strength, demand extra validation data/ TEE, autoencoders, AEs/watermarks/blockchain overheads, sensitivity to non-IID, and hyperparameter tuning. |
| Preserving global model integrity | ( Sultana et al., 2022 ) , ( Liu et al., 2022b ) , ( Du et al., 2023 ) , ( Cao et al., 2022 ) , ( Uprety and Rawat, 2021 ) , ( Andreina et al., 2021 ) , ( Issa et al., 2024 ) | Global integrity assessment across training rounds, decentralized multi-model aggregation with voting, trust scores via global / client validation, and privacy-aware and personalized encoders. | Post-aggregation checking, detect and neutralize persistent poisoning effects, majority voting across models, maintain privacy while preserving utility. | Global accuracy under strong poisoning; integrity or trust scores; fraction of compromised updates detected and discarded; privacy-utility trade-off metrics. | EMNIST, FEMNIST, MNIST, CIFAR-10, synthetic datasets. | Adds a global sanity check layer beyond local defenses, can recover from polluted rounds, and reduces a single point of failure via multiple globals. | Depends on representative validation sets/clients, higher coordination/computation overhead, and personalized encoders have extra complexity. |
| Against malicious servers | ( Le et al., 2023 ) , ( Mothukuri et al., 2021 ) , ( Jeter et al., 2023 ) , ( Huang et al., 2024 ) , ( Han et al., 2024 ) , ( Hao et al., 2023 ) , ( Guo et al., 2023 ) , ( Zhou et al., 2020 ) , ( Ma et al., 2022a ) , ( Ye et al., 2022 ) , ( Liu and Shang, 2022 ) , ( Tang et al., 2024 ) | Server-side threat analyses, image augmentation, enhanced ciphertext schemes, decentralized / blockchain FL, auditing and monitoring, limit or detect client identification by the server. | Limit gradient inversion, server-side reconstruction, client identification, protect against server model manipulation, and accountability for server operations using decentralization and auditing. | Privacy leakage metrics (e.g., inversion success), identification success, global accuracy under a malicious server, audit, and consensus performance. | CIFAR-10/100, MNIST, synthetic datasets. | Targets the strongest adversary (the server) explicitly, preserves utility, and increases transparency and accountability. | Blockchain/auditing introduces latency, storage, and system complexity, tailored to specific modalities (e.g., images), and cryptography demands careful key and protocol management. |
| Adversarial-resilient aggregation | Refer section 3.2 | Statistical robust aggregation, byzantine-resilient aggregation, similarity/distance-based methods, reputation-based methods, cryptography-enhanced methods | Tolerate adversarial or corrupted gradients, detect and suppress abnormal updates and outliers during aggregation, account for reliable participants, protect update confidentiality and integrity. | Global accuracy under attack, ASR, Byzantine tolerance level, convergence stability, computation and communication overhead. | MNIST, CIFAR-10/100, EMNIST, FEMNIST ImageNet | Enhances robustness against adversarial behaviors, tolerates a fraction of malicious clients, enhances stability and reliability of global model updates. | Assumes bounded adversaries, performance may get affected under highly non-IID data, similarity and reputation mechanisms introduce additional computation, cryptography-enhanced methods increase overhead. |
| Advanced cryptographic strategies | ( Xia et al., 2023 ) , ( Sotthiwat et al., 2021 ) , ( Zhu et al., 2021 ) , ( Ma et al., 2024 ) , ( Javeed et al., 2024 ) , ( Chehimi and Saad, 2022 ) , ( Gurung et al., 2023 ) , ( Ma et al., 2022d ) , ( Moriai, 2019 ) , ( Moshawrab et al., 2023 ) | HE, partial HE, two-trapdoor HE with secure cosine similarity, ZKPs for verifiable training, PQC, quantum-resistant signatures + blockchain, PE for SVM-based FL. | Guarantee confidentiality and integrity of updates and aggregation, public verifiability, and quantum-resistant security. | Byzantine tolerance level, runtime and memory overhead, proof generation/verification time/size, comm. cost, task accuracy under encrypted/PQC/PE. | MNIST, CIFAR-10, synthetic datasets. | Offer strong, provable security guarantees, can be combined with other defenses, higher tolerance in heterogeneous environments. | Higher overhead, especially for deep models, polynomial approximations needed for deep non-linear models, complex key management, large-scale deployment challenges. |
| Insight | Core idea | Defended threats | Typical design pattern | Advantages | Drawbacks |
| Conventional defenses fail in FL | Classical centralized defenses that assume direct access to training data and full control over the pipeline conflict with FL’s distributed approach. | Data poisoning in raw data, data-level outliers, and direct anomaly detection at the sample level. | Assume the server sees the full dataset; reweight or discard suspicious samples using robust losses or feature statistics. | Well-understood in centralized ML; strong guarantees only when data is pooled. | Incompatible with FL’s update-only interface, violates privacy constraints, and cannot inspect raw local data or intermediate features. |
| Multi-attack coverage | Defenses that simultaneously address multiple threats appear more appealing than mechanisms tailored to a single attack type, motivating taxonomies organized by phase or location rather than by attack. | Label-flipping, gradient poisoning, backdoors, inference or reconstruction attacks. | A combination of robust aggregation with DP noise, KD, autoencoder-based update scoring, or sparsification. | One mechanism spans multiple stages (local, aggregation, global) and minimizes attack-specific tuning, making it convenient in mixed-threat environments. | Difficult to ensure the best protection for any single attack, covert and adaptive adversaries can still circumvent generic rules, DP noise, and robustness interaction is non-trivial. |
| Layered client-server protection | Combined protection across local training, aggregation, and system layer for increased effectiveness than any isolated component. | Local poisoning; gradient leakage; compromised servers; integrity of global model and logs. | Client (preprocessing, DP, VAE, KD, etc.), server (similarity or trust scoring, robust aggregation), system (blockchain, auditing, watermarking). | Suppress attacks early at the client, preserve long-term traceability and accountability of updates, and model lineage. | Increased overall complexity, difficulty in coordinating thresholds and hyperparameters across layers, and overhead on resource-constrained devices. |
| Cryptography boosts robustness | Strong cryptography enables confidentiality and verifiability, but is resource-intensive. | Inference attacks, model or gradient inversion, leakage, poisoning, and replay attacks. | Secure aggregation, ZKP, PE, and PQC techniques | Securely fetch individual updates from the server and peers, allow validation during aggregation and training, tolerate stronger threat models (e.g., curious or semi-malicious servers). | Encrypted computation is slow (high runtime), nonlinear networks require polynomial approximations, and protocols are complex and involve key management issues. |
| Fragile defensive assumptions | Simple assumptions that may break under sophisticated or more realistic conditions. | High-Byzantine poisoning, adaptive and covert adversaries that target defense logic, highly non-IID clients. | Assume a maximum fraction of suspicious clients, need good validation data or honest clients, threshold tuning for specific non-IID levels. | Offer provable bounds or strong empirical robustness under an assumed context, easy to analyze and implement for targeted scenarios. | Bound on malicious fraction; need auxiliary data or TEE; degrade under extreme non-IID. |
| Robustness through personalization | Client or cluster-specific models to reduce the impact of poisoned or biased updates. | Client-targeted poisoning; non-IID amplified attacks; fairness/bias issues. | Clustered aggregation, personalized layers, trust-weighted personalization, KD-based refinement, representation or prototype-level adaptation. | Reduces cross-client poisoning; isolates malicious clusters, better fits local data, improves accuracy and robustness on non-IID data. | More models to manage, difficulty in global guarantees, and misassignment of clusters strengthen attacks. |
| Dataset | Type | Description | IID | Non-IID | Applied Scenario |
| CIFAR-10 | Image Classification | 60,000 32x32 color images across 10 classes | Object recognition, autonomous systems. | ||
| CIFAR-100 | Image Classification | 60,000 32×32 color images across 100 classes | Fine-grained classification, FL model evaluation purposes. | ||
| MNIST | Handwritten Digits | 70,000 28×28 grayscale images (digits 0-9) | Handwritten digit recognition, basic FL benchmarking. | ||
| Extended MNIST | Handwritten Characters | Includes handwritten letters and digits | Real-world non-IID settings for character recognition tasks. | ||
| Fashion-MNIST | Fashion Image Classification | 70,000 28×28 grayscale images across 10 fashion categories | Retail analytics, product categorization. | ||
| Shakespeare | Text (NLP) | Complete works of Shakespeare divided by characters | Personalized language modeling, and text generation tasks. |
| Focus Area | Key Robust Strategies |
|---|---|
| Strengthening Privacy and Resilience | - Differential Privacy with noise injection - Homomorphic and polymorphic encryption - Post-Quantum Cryptography (PQC) - Blockchain and SMC for decentralized trust |
| Designing Scalable Federated Architectures | - Lightweight learners (e.g., ELM) for edge nodes - Hierarchical, cluster-based FL frameworks - Decentralized aggregation to mitigate central bottlenecks - Quantum ML and Hybrid Classical-Quantum Systems - Asynchronous FL to control latency and device churn. |
| Advancing Personalization and Adaptivity | - Personalized models using meta-learning and local fine-tuning. - Use of adaptive optimizers for faster convergence - Multi-task learning for promoting personalized models. - Client-specific loss functions or model layers to support heterogeneity - Semi-supervised and multimodal FL approaches |
| Enhancing Robustness and Model Integrity | - Robust aggregators to deal with malevolent or corrupted updates - Reputation-driven solutions for client selection - Quantum-driven optimization - Fault tolerance against participant dropouts and stragglers - Dynamic selection of clients based on device availability and reliability. |
| Robust Communication Efficiency | - Model pruning, gradient compression to reduce overhead - Sparsity and quantization for efficient aggregation - Asynchronous and layer-wise updates - Communication scheduling based on device capacity |