Research
MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
Overview Research area: Distributed machine learning, specifically peer-to-peer federated learning (P2P FL) for wireless and edge networks. Technical level: Advanced (assumes familiarity with federate

- arXiv
- 2512.05234
- Published
- 2025-12-04
- Authors
- Felix Mulitze, Herbert Woisetschläger, Hans Arno Jacobsen
AI summary
Overview
- Research area: Distributed machine learning, specifically peer-to-peer federated learning (P2P FL) for wireless and edge networks.
- Technical level: Advanced (assumes familiarity with federated learning, communication complexity, knowledge distillation, and differential privacy).
- Scope: The paper proposes and empirically evaluates MAR-FL, a serverless P2P federated learning system that uses iterative group-based aggregation to cut communication cost from O(N²) to O(N log N) while tolerating peer churn.
What This Paper Is About
Peer-to-peer federated learning removes the central server bottleneck, but existing P2P methods pay for it with very high communication cost — every peer effectively talks to many others. MAR-FL attacks that cost directly: peers repeatedly form small, changing groups and average models only within those groups, so information still spreads globally without all-to-all exchange. The paper's goal is to show that this group-based scheme keeps model quality equal to standard baselines while using far less communication and remaining robust when peers drop out.
Key Contributions
- A P2P FL system with reduced communication complexity. MAR-FL achieves O(N log N) communication per iteration, versus the O(N²) complexity of all-to-all P2P FL, by adapting Moshpit All-Reduce (MAR) group-based aggregation to federated learning.
- Integration of Knowledge Distillation for convergence acceleration. The authors introduce Moshpit-KD (MKD), which selects top-ℓ teachers by lowest KL divergence and distills from them in the first K FL iterations, with a linearly decaying weighting of the KL loss term.
- Fully decentralized differential privacy. They adapt DP-FedAvg with adaptive clipping to a serverless setting, so that privacy loss accrues from local computations while MAR only averages already-privatized models.
- A comprehensive experimental comparison. MAR-FL is benchmarked against client-server FedAvg, RDFL (Galaxy's ring all-reduce), and a naive all-to-all AR-FL baseline on MNIST and 20NG, including tests of scalability, partial participation, dropout, and DP.
Main Findings
- Communication efficiency up to 10× better: Across both MNIST and 20NG, MAR-FL matches the training performance of RDFL and AR-FL while requiring up to 10 times less communication per iteration.
- Identical model utility under suitable parameters: With group size 5 and 3 MAR rounds for 125 peers (since 125 = 5³), each MAR-FL iteration attains an exact global average, so parity with baselines is expected.
- MKD further reduces total communication: With MKD, MAR-FL needs over 2 times less communication to reach 50% accuracy on 20NG, although the per-iteration communication load increases; the trade-off is tunable via the number of KD iterations.
- Partial participation hurts, churn does not: Partial participation causes substantial degradation in training performance, while configured network churn and unreliable connectivity cause no additional accuracy drops — a pattern shared by all three baselines.
- Robust net advantage under disturbance: Even with 50% participation and 20% dropout likelihood, RDFL and AR-FL require more than 5 times the communication of MAR-FL to reach the same model utility.
- DP behaves like standard FedAvg: Raising the DP noise multiplier σ reduces privacy loss ε but eventually degrades model utility, matching observations for FedAvg with DP.
- Convergence bound: The expected average distortion after T averaging iterations decays as ((r−1)/N + r/N²)^T, and this bound is independent of the spectral properties of the communication graph, unlike gossip-based decentralized FL.
- Remaining gap to client-server FL: The paper states that a performance gap toward client-server FL still exists, which it lists as a limitation.
Methodology in Plain English
Each of the N peers holds a private, possibly non-i.i.d. local dataset. In every FL iteration, participating peers perform local Momentum-SGD updates on a fixed number of mini-batches. Then a coordination step — handled through a Hivemind Kademlia distributed hash table that carries only lightweight barriers and group-formation metadata, never model weights — partitions peers into small groups. Peers average models and momentum vectors only within their group, and this group formation repeats over multiple rounds. Peers avoid meeting the same partner twice within one FL iteration by deriving their group key from their previous round's chunk index. A single DHT lookup costs at most O(log N) hops; because the implementation occasionally scans peer announcements with O(N) look-ups, the control-plane cost per round is O(N log N).
Optionally, MKD is applied in the first K iterations: peers collect candidate teacher models through the same group-formation procedure, rank them by KL divergence between softened output distributions, pick the top ratio (ρℓ = 0.4), and train as students on a weighted sum of a temperature-scaled KL loss (τ = 3.0) and cross-entropy on hard labels, with the KL weight shrinking linearly. For privacy, each peer clips its local model delta to an adaptive bound, adds Gaussian noise, smooths the privatized delta with factor β = 0.9, and runs MAR on the privatized models; the clipping bound is then updated from a globally averaged clipping rate.
Experiments use MNIST with a CNN and 20 Newsgroups with a frozen DistilBERT plus classification head, split non-i.i.d. via Latent Dirichlet Allocation (α = 1.0) across 16, 64, and 125 peers, with 64 samples per MNIST peer and 16 per 20NG peer per aggregation round. Training uses SGD with momentum (η = 0.1, μ = 0.9), full participation by default, and one node with 4 H100 GPUs, 768 GB memory, and 96 CPU cores. Because of simulation constraints, model evaluation happens every fifth FL iteration. BrainTorrent and SAPS are discussed but excluded as baselines due to their communication limitations; Galaxy Federated Learning as a whole is not compared because of its reliance on a distributed ledger for verification.
Why This Matters
Impact on research. The paper shows that the O(N²) communication cost widely assumed to be inherent to fully decentralized aggregation can be reduced to O(N log N) without losing model quality, and that a convergence bound free of communication-graph spectral properties is achievable. It also gives a concrete recipe for combining serverless aggregation with knowledge distillation and with DP adaptive clipping, both of which are normally formulated around a central server.
Real-world applications.
- Wireless and 6G/Wi-Fi 9 edge networks, where peers are bandwidth-limited and connectivity is unreliable.
- Multi-operator collaborations or community-driven deployments where no single entity can or should control the training process.
- Cross-silo FL across organizations that cannot move data across geographical or regulatory boundaries.
- Regions with limited power grid capacity or infrastructure budgets that cannot build large centralized AI data centers.
Industry relevance. The paper frames communication cost as an economic feasibility issue: RDFL's cost is described as orders of magnitude higher than centralized FedAvg, making it economically infeasible for wireless environments. By narrowing the gap to FedAvg while eliminating the central server, MAR-FL targets settings where server-side compute, memory, or networking capacity is a bottleneck or a single point of failure.
Future Directions
- A thorough analysis of partial participation and network churn to bring the system closer to real-world applicability.
- Exploring approximate aggregation and adaptive group-based information propagation to further improve communication efficiency and narrow the gap to client-server FedAvg.
- Experimental evaluation of DP in MAR-FL that exploits the system's scalability to compress peer-sampling rates, maintaining model utility while reducing privacy loss.
- Analysis of how group-based aggregation combined with momentum affects DP dynamics, which the authors state remains open.
Target Audience
Researchers and practitioners working on decentralized and federated machine learning, edge and wireless AI systems, and distributed training infrastructure. It is most useful to readers already comfortable with FL aggregation schemes, communication complexity analysis, knowledge distillation, and differential privacy, and to engineers evaluating whether P2P FL is deployable under bandwidth and churn constraints.
Authors’ abstract
The convergence of next-generation wireless systems and distributed Machine Learning (ML) demands Federated Learning (FL) methods that remain efficient and robust with wireless connected peers and under network churn. Peer-to-peer (P2P) FL removes the bottleneck of a central coordinator, but existing approaches suffer from excessive communication complexity, limiting their scalability in practice. We introduce MAR-FL, a novel P2P FL system that leverages iterative group-based aggregation to substantially reduce communication overhead while retaining resilience to churn. MAR-FL achieves communication costs that scale as O(N log N), contrasting with the O(N^2) complexity of previously existing baselines, and thereby maintains effectiveness especially as the number of peers in an aggregation round grows. The system is robust towards unreliable FL clients and can integrate private computing.