LFHE: Local-First Heuristic Evolution for Bounded Local Topology Search in Decentralized Learning with Non-IID Data
Organizations: Durham University · University of Cambridge
Abstract
Decentralized learning is highly sensitive to communication topology under non-IID data. Adaptive peer-selection methods can exploit local model information, but broader peer discovery may require increasingly large control state, whereas direct spectral optimization typically relies on graph-wide information. We study the intermediate setting of bounded local topology search and propose Local-First Heuristic Evolution (LFHE), a representation-driven rewiring framework whose candidate discovery and scoring use only ego-neighborhood and friend-of-a-friend (FoF) information. The structural score admits an exact interpretation through graph Dirichlet energy: its sum across clients equals twice the representation Dirichlet energy, which under standard linear consensus dynamics governs the instantaneous dissipation of representation disagreement. LFHE combines this state-dependent structural signal with early exploration and degree control, while algebraic connectivity remains an offline graph diagnostic. Under bounded sparse degree, its FoF candidate state remains local rather than expanding toward population-wide peer tracking. Across four image, speech, and text benchmarks, LFHE achieves competitive decentralized learning performance. Matched-protocol controls identify the structural term as the principal empirical topology-selection signal, while comparison with broader peer discovery exposes a trade-off between predictive performance and discovery-state locality. Together, these results motivate state-aware bounded local topology search between pairwise peer selection and globally informed topology optimization.
Figures & tables
| Method | CIFAR-10 ( 70%) | CIFAR-100 ( 35%) | Speech Commands ( 70%) | Sentiment140 ( 65%) | ||||
|---|---|---|---|---|---|---|---|---|
| R@T | Final Acc. (%) | R@T | Final Acc. (%) | R@T | Final Acc. (%) | R@T | Final Acc. (%) | |
| FedAvg | 115 | 78.40 0.80 | 85 | 51.90 0.30 | 35 | 86.80 0.50 | 40 | 70.60 2.70 |
| Fully Connected | 105 | 78.40 0.70 | 85 | 51.80 0.50 | 35 | 86.40 1.40 | 40 | 70.60 2.60 |
| Ring | – | 52.60 1.20 | – | 27.00 0.80 | – | 53.00 3.10 | – | 58.00 2.70 |
| Static Random | – | 67.40 1.70 | 160 | 41.70 1.20 | 75 | 74.20 3.50 | – | 63.50 2.30 |
| Epidemic Learning | – | 70.00 1.60 | 130 | 44.10 0.70 | 60 | 78.10 3.30 | 190 | 65.50 2.30 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Connectivity-aware | Similarity-aware | Ego-graph local | No global graph statistics |
|---|---|---|---|---|
| Spectral methods | ✓ | |||
| D-PSGD / Gossip | ✓ | |||
| D-Cliques / STL-FW | ✓ | ✓ | ||
| Epidemic Learning | ✓ | ✓ | ||
| DissDL (legacy impl.) | ✓ | ✓ | ✓ | |
| Morph | ✓ | ✓ |
| Item | Setting |
|---|---|
| Number of clients | for the main experiments; for scalability experiments. |
| Online degree threshold | ; the expected-degree initializer is not guaranteed to satisfy a hard cap. |
| Topology update interval | communication rounds. |
| Evaluation interval | Every 5 communication rounds. |
| Random seeds | . |
| Non-IID partitioning | Class-wise Dirichlet partitioning with by default; the component ablation and moderate-scale diagnostic use . |
| Dataset | Model architecture | Representation |
|---|---|---|
| CIFAR-10 | CNN with two convolutional blocks: Conv(3,32)-BN-ReLU, Conv(32,32)-BN-ReLU, MaxPool; Conv(32,64)-BN-ReLU, Conv(64,64)-BN-ReLU, MaxPool; adaptive average pooling to ; MLP classifier with dropout 0.3. | Flattened final classifier-layer weights. |
| CIFAR-100 | Same CNN backbone as CIFAR-10, with the final classifier dimension changed to 100 classes: . | Mean-reduced final classifier-layer weights, producing a 100-dimensional representation. |
| Speech Commands | CNN over log-Mel spectrograms: Conv(1,32)-BN-ReLU, Conv(32,32)-BN-ReLU, MaxPool; Conv(32,64)-BN-ReLU, Conv(64,64)-BN-ReLU, MaxPool; adaptive average pooling to ; MLP classifier with dropout 0.3. | Flattened final classifier-layer weights. |
| Sentiment140 | Embedding-based text classifier with 128-dimensional token embeddings, attention-mask mean pooling, and MLP classifier with dropout 0.2. | Flattened final classifier-layer weights. |
| Method | Final accuracy (%) | nAUC (%) |
|---|---|---|
| LFHE | 71.37 2.19 | 60.57 1.52 |
| Morph | 72.25 1.92 | 63.17 1.23 |
| Order | Final acc. (%) | nAUC (%) | R@65 | Final | Edge Jaccard |
|---|---|---|---|---|---|
| Fixed | 69.03 | 58.82 | 160 | 0.936 | 1.000 |
| Reverse | 70.54 | 60.29 | 145 | 0.685 | 0.291 |
| Random/update | 69.61 | 59.24 | 155 | 0.807 | 0.271 |
| Method | Final | Final Acc. | Mean Div. | Mean Rewires | |
|---|---|---|---|---|---|
| Static | 0.592 | 0.000 | 0.6675 | 0.766 | 0.000 |
| Similarity-only | 0.578 | -0.014 | 0.6704 | 0.727 | 0.183 |
| Exploration-only | 1.038 | 0.447 | 0.7356 | 0.221 | 3.406 |
| Structural-only | 2.306 | 1.715 | 0.7508 | 0.136 | 1.111 |
| LFHE | 2.247 | 1.655 | 0.7482 | 0.133 | 1.333 |