cs.LGAug 16, 2025

FedUHD: Unsupervised Federated Learning using In-Memory Hyperdimensional Computing

Authors: You Hak Lee, Keming Fan, Xiaofan Yu, Quanling Zhao, Tianqi Zhang, Flavio Ponzina, Tajana Rosing

Organizations: University of California San Diego, USA

Abstract

Unsupervised federated learning (UFL) enables privacy-preserving distributed training without data labeling, yet practical deployment remains challenging due to non-IID data, high computational and communication costs at edge devices, and sensitivity to communication noise. We propose FedUHD, the first UFL framework based on Hyperdimensional Computing (HDC). On the client side, FedUHD employs kNN-based cluster hypervector removal to mitigate non-IID effects by filtering detrimental local outliers. On the server side, cluster-aware HDC aggregation leverages cluster-level statistics to stabilize learning across heterogeneous clients. To further improve efficiency, we design a compute-in-memory (CIM) accelerator based on a novel phase-change memory (PCM) device, integrated with lightweight ASIC digital modules to execute the client-side HDC pipeline within the accelerator. The intrinsic robustness of HDC to low precision and device variations enables efficient mapping onto analog PCM crossbars, exploiting massive parallelism while minimizing data movement. Experimental results show that FedUHD achieves comparable accuracy to state-of-the-art neural network-based UFL methods across all datasets. On HAR and CIFAR10/100, FedUHD delivers an average 2,239x speedup and 1,542x higher energy efficiency on GPU. In addition, FedUHD reduces communication cost by up to 176x on HAR and CIFAR10/100 and demonstrates greater robustness than Orchestra under communication noise. Compared to GPU implementation of FedUHD, the proposed PCM-based accelerator provides an additional 4.07x speedup and three orders of magnitude higher energy efficiency on average. Furthermore, the results demonstrate the benefit of PCM over RRAM as a CIM substrate.

Figures & tables

Explore similar work

May 12, 2026cs.LG

Fed-BAC: Federated Bandit-Guided Additive Clustering in Hierarchical Federated Learning

Hierarchical federated learning (HFL) leverages edge servers for partial aggregation in edge computing. Yet existing FL methods lack mechanisms for jointly optimizing cluster assignment and client selection under data heterogeneity. This paper proposes Fed-BAC, which integrates additive cluster personalization with a two-level bandit framework: contextual bandits at the cloud learn server-to-cluster assignments, while Thompson Sampling at each edge server identifies high-contributing clients. The additive decomposition enables the sharing of knowledge between groups through a globally aggregated network, while cluster-specific networks capture distribution variations. Across three classification benchmarks (CIFAR-10, SVHN, Fashion-MNIST) under moderate (α=0.5α= 0.5) and severe (α=0.1α= 0.1) Dirichlet non-IID partitioning, Fed-BAC achieves distributed accuracy gains of up to +35.5pp over HierFAVG and +8.4pp over IFCA, while requiring only 80% client participation, converging 1.5 to 4.8×\times faster depending on dataset and accuracy target, and improving cross-server fairness. These gains are further validated at 5×\times deployment scale on CIFAR-10. The advantage of Fed-BAC increases with heterogeneity severity, confirming that additive cluster personalization becomes increasingly valuable as data distributions diverge.
Jun 9, 2026cs.LG

Accurate and Resource-Efficient Federated Continual Learning

Federated continual learning (FCL) must learn from distributed task streams under limited resources, such as communication, computation, memory, and label availability. Existing FCL methods often rely on repeated local optimization, replay, and full supervision. Analytic alternatives avoid iterative training and replay, but using high-dimensional random features to improve accuracy requires a second-order feature statistic, the Gram matrix, which has a quadratic communication cost in the random feature size MM. We propose FedRAN, a resource-aware analytic FCL framework that replaces gradient-based updates with compact random feature statistics. Each client transmits a truncated-SVD summary of its Gram matrix, reducing the dominant second-order upload from quadratic to linear in MM for fixed rank. The server performs a two-level QR-SVD subspace merge, spatially across clients and temporally across tasks, and solves a ridge classifier in closed form. FedRAN further supports label scarcity through prototype-based pseudo-labeling. Across CIFAR-100, ImageNet-R, and VTAB datasets, FedRAN improves average accuracy by up to 4.8 percentage points over the strongest baseline, uses 30.6-121.8×\times less per-client communication than optimization-based FCL, and is 190.3×\times faster on average than gradient-based baselines; with only 20% labels, pseudo-labeling improves average accuracy by up to 6.61 points. These results show that FedRAN enables accurate and resource-efficient FCL under communication, computation, and label constraints. The source code is available at https://github.com/JebacyrilArockiaraj/Fed-RAN-SSL.
Aug 10, 2026cs.LG

FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients

Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per deployment limit is costly. Federated supernet training instead learns one elastic model with differently sized subnetworks, then deploys a suitable one to each device. When client inference budgets differ, however, parameters exclusive to high-cost subnetworks are reachable by fewer clients. We propose FEAST, a federated shared-space training framework that counters this imbalance by jointly training multiple subnetworks within each client's limit. Budget-tailored sub-supernet routing sends only the relevant supernet portion, and sparse aggregation merges the returned parameter slices. The trained supernet directly serves the subnetworks used during federation and supports post-hoc extraction of additional subnetworks without federated retraining. We further show that independently assigning clients' training-data volumes and inference budgets can distort accuracy--inference-cost comparisons in heterogeneous FL simulations, and introduce a one-parameter γγ-allocation protocol to control this coupling. In our experimental setup, the SuperFedNAS and DeepFedNAS supernet training procedures remain near chance at 25M and reach at most 17.09%17.09\% at 596596M inference MACs; FEAST reaches 71.06%71.06\% at 596596M, 2.42.4 points above the strongest model-heterogeneous weight-sharing baseline at its largest tier. Across CIFAR-100, CINIC-10, and TinyImageNet-200, FEAST achieves the highest population-averaged accuracy among the evaluated weight-sharing methods when each client receives its largest affordable subnetwork. Sub-supernet routing reduces aggregate model-parameter traffic by 6.8×6.8\times relative to full-supernet transmission.