cs.LGMay 26, 2025

Federated Independent Component Analysis via Spectral Alignment and Robust Aggregation

Authors: Xin Bing, Dian Jin, Yuqian Zhang

Organizations: University of Toronto · National University of Singapore · Rutgers University

Abstract

This paper studies robust estimation for Independent Component Analysis (ICA) in the federated learning setting, where data are locally distributed across clients and may exhibit substantial heterogeneity. The goal is to recover a common global mixing matrix by aggregating local estimators computed by individual clients. The main difficulty is that local ICA estimators are identifiable only up to sign flips and column permutations and may have highly heterogeneous estimation quality. We propose a novel three-step aggregation method that first aligns the permutations across all local estimators via a particular spectral clustering approach, then aligns the signs within each estimated cluster via another tailored spectral approach, and finally applies the geometric median for robust aggregation. The proposed estimator is shown to remain accurate even when a substantial fraction of local estimators are of low quality or inconsistent, as long as each cluster contains a majority of accurate estimators. This contrasts with its analogue based on simple averaging, whose performance is determined by the worst local estimator. In the homogeneous setting, the robust estimator is also shown to be minimax optimal, up to a logarithmic factor, both when the number of clients remains fixed and when it diverges. The theoretical findings are corroborated by simulation studies and a real-data analysis.

Figures & tables

Explore similar work

Sep 17, 2026cs.LG

Distributionally Robust Federated Learning with Multi-Source Data

Federated learning trains a shared model from private client data. In practice, data-generating distributions may differ, and the true mixture across clients is often unknown, making the underlying group distribution difficult to specify. Existing approaches address cross-client mixture uncertainty by optimizing against the worst-case mixture, yet assume accurate client-wise distribution estimates. However, these estimates can be unreliable when based on finite samples. To handle both cross-client mixture uncertainty and within-client distributional ambiguity, we construct a global ambiguity set as the union of admissible mixtures of local ambiguity sets. The construction allows client-specific ambiguity radii and admits a client-wise separable reformulation. Leveraging this structure, we establish a high-probability out-of-sample performance guarantee. We further develop a federated algorithm for a penalty-based reformulation and prove its convergence under milder regularity conditions. Simulations validate the algorithm's effectiveness.
Apr 24, 2026cs.LG

Data-Free Contribution Estimation in Federated Learning using Gradient von Neumann Entropy

Client contribution estimation in Federated Learning is necessary for identifying clients' importance and for providing fair rewards. Current methods often rely on server-side validation data or self-reported client information, which can compromise privacy or be susceptible to manipulation. We introduce a data-free signal based on the matrix von Neumann (spectral) entropy of the final-layer updates, which measures the diversity of the information contributed. We instantiate two practical schemes: (i) SpectralFed, which uses normalized entropy as aggregation weights, and (ii) SpectralFuse, which fuses entropy with class-specific alignment via a rank-adaptive Kalman filter for per-round stability. Across CIFAR-10/100 and the naturally partitioned FEMNIST and FedISIC benchmarks, entropy-derived scores show a consistently high correlation with standalone client accuracy under diverse non-IID regimes - without validation data or client metadata. We compare our results with data-free contribution estimation baselines and show that spectral entropy serves as a useful indicator of client contribution.
Date pendingcs.LG

Class-wise Contribution Estimation via Logit Maximization for Federated Learning

Federated learning (FL) enables collaborative learning of computer vision models, where privacy and regulatory constraints prevent centralizing data across devices or organizations. However, practical FL deployments often exhibit severe class imbalance and label skew, causing standard aggregation protocols to overfit dominant clients and degrade minority-class performance. We propose a data-free, class-wise contribution estimation and aggregation framework based on logit maximization (CELM) that does not require sharing raw data, client metadata, or auxiliary public datasets. The FL server probes client updates to obtain class-wise evidence scores and assembles a cross-client evidence matrix, which quantifies both per-class competence and class coverage. Using this matrix, we compute contribution weights that upweight clients providing strong, discriminative evidence for underrepresented classes. The resulting aggregation is stable due to simplex constraints and momentum smoothing, and it remains compatible with standard FL training pipelines. We evaluate the approach on representative vision benchmarks under controlled non-IID and pathological label splits, demonstrating that CELM-based aggregation improves robustness to imbalance and statistical heterogeneity, while yielding better performance without requiring any additional data exchange.