Research
Energy and Memory-Efficient Federated Learning With Ordered Layer Freezing
Energy and Memory-Efficient Federated Learning With Ordered Layer Freezing Overview Research area: Federated learning (FL) for the Internet of Things (IoT), specifically resource-efficient training on

- arXiv
- 2512.23200
- Published
- 2025-12-29
- Authors
- Ziru Niu, Hai Dong, A. K. Qin, Tao Gu, Pengcheng Zhang
AI summary
Energy and Memory-Efficient Federated Learning With Ordered Layer FreezingOverview
Research area: Federated learning (FL) for the Internet of Things (IoT), specifically resource-efficient training on edge devices with heterogeneous compute, memory, and bandwidth.
Technical level: Intermediate. The core idea (freezing low-level layers and training only top layers) is intuitive, but the paper also includes a non-convex convergence analysis with formal assumptions and theorems, which is more advanced.
Scope (1 sentence): The paper proposes FedOLF, a federated learning framework that consistently freezes low-level layers in a predefined order on resource-constrained clients and combines this with Tensor Operation Approximation (TOA) to reduce computation, memory, and communication costs while preserving accuracy on non-iid client data.
What This Paper Is About
Federated learning lets many edge devices train a shared model without sending their raw data to a server, but IoT devices often lack the processor, memory, and bandwidth needed to train and transmit deep neural networks. Existing fixes such as dropout (training a pruned sub-model) and layer freezing (freezing some layers while keeping the full architecture) either hurt accuracy on non-iid data or quietly consume large amounts of memory during backpropagation. FedOLF addresses both problems by always freezing the lowest layers rather than a random or top-first subset, and by compressing only those frozen layers for transmission.
Key Contributions
-
An analysis of why prior layer-freezing methods fail on memory. The authors show that random layer freezing (as in CoCoFL) and top-first freezing (as in SLT) still require large memory because the top frozen layers must store activations and gradients to propagate error back to low-level active layers. Ordered (bottom-first) freezing gives a shorter backpropagation path and lower memory use.
-
The FedOLF framework. Resource-constrained clients consistently freeze low-level layers and train only the remaining top-level layers, with layer counts chosen per device capacity. The server decomposes the global model into frozen and active sets, clients train the active layers with SGD, and only the updated active layers are uploaded. Aggregation uses the same layer-wise scheme as prior work (CoCoFL/SLT lineage).
-
Combination with Tensor Operation Approximation (TOA). The server sparsifies the frozen layers before download using weighted sampling of tensors (filters or neurons) with probabilities proportional to their Frobenius norms, where the number of retained tensors is
floor(s * H_q)andsis a scaling factor in(0, 1]. TOA is applied to all frozen layers except the last one, so the representation dimensions feeding the active layers stay unchanged. The paper states only the frozen layers are approximated, unlike the original TOA, to ensure all active layers are fully trained. The authors state this reduces downstream communication cost by approximatelyO(s^2)and claim this paper is among the first to apply TOA to reduce FL memory and communication costs. -
A convergence analysis for non-convex settings, plus an empirical evaluation across EMNIST (CNN), CIFAR-10 (AlexNet), CIFAR-100 (ResNet20 and ResNet44), and CINIC-10 (ResNet20 and ResNet44) under non-iid client data.
Main Findings
-
Accuracy over non-iid data: The paper reports that FedOLF achieves at least 0.3%, 6.4%, 5.81%, 4.4%, 6.27% and 1.29% higher accuracy than existing works on EMNIST (with CNN), CIFAR-10 (with AlexNet), CIFAR-100 (with ResNet20 and ResNet44), and CINIC-10 (with ResNet20 and ResNet44) respectively, together with higher energy efficiency and lower memory footprint.
-
Ordered freezing reduces memory, random freezing does not: Using ResNet20 on CIFAR-100 and measuring required memory with PyTorch's
TORCH.CUDA.MAX_MEMORY_ALLOCATED, the authors report that random layer freezing does not effectively relieve memory requirements, because oversized activation maps must still be stored across all frozen layers for backpropagation. Ordered layer freezing notably mitigates required memory by shortening the backpropagation path. The paper's figures show this comparison but the summary content does not report the specific measured memory values. -
Gradient error from freezing is bounded: Because frozen layers are stale, they produce a representation that diverges from the true one (
x'_{l_k} = x_{l_k} + σ_{l_k}), and the resulting approximate gradient∇f'_kdiffers from∇f_k. Drawing on prior work, the paper argues the representation error and gradient error are usually upper bounded, which limits the accuracy damage from freezing. -
Low-level layers are largely redundant across clients: Citing prior work on Centered Kernel Alignment (CKA), the paper states that low-level layers across local models tend to have high CKA similarity, meaning a resource-constrained client can "borrow" these low-level layers from the server (trained by more powerful clients in earlier rounds) with limited error.
-
Convergence behavior: Under assumptions of
L-smoothness, bounded local-gradient varianceγ, and bounded layer-freezing gradient divergenceD, the analysis shows the objective decreases until reaching anε-critical point inT = O(‖w^ε − w^0‖ / (εη))iterations. Two regimes are given: whenη ≤ 1/L, the critical-point bound isε = ε₂ = D + γ; when1/L < η < 3/(2L), the bound isε = ε₁ = (D(ηL − 1) + sqrt(ηD²L + 8ηLγ² + 6ηDLγ + D² − 3γ²)) / (3 − 2ηL). -
Frozen layers are cheaper in both compute and bandwidth: Clients download the full active layers plus a sparsified version of the frozen layers, then upload only the updated active layers, so frozen-layer gradients, activations, and transmissions are avoided.
Methodology in Plain English
The authors begin by diagnosing a problem with existing layer-freezing schemes: if you freeze layers at the top of a network and keep low layers trainable, error signals travelling backward still pass through the frozen top layers, forcing the system to keep their activations and gradients in memory. They verified this empirically by implementing both random and ordered freezing on ResNet20 with CIFAR-100 and reading off PyTorch's peak allocated GPU memory.
Their fix flips the direction: always freeze the low-level layers and keep the top layers trainable. Since the lowest frozen layers do not feed into the gradient computation of the layers above them, they store nothing during backpropagation, so the backward path is short and memory use drops. Each client's number of frozen layers is set by its device capacity; a powerful client can freeze zero layers and train the whole model.
On top of this, the server compresses the frozen layers before sending them. For each frozen layer except the last, it keeps floor(s * H_q) tensors chosen by weighted sampling, where the sampling probability of each tensor is proportional to its Frobenius norm — a choice the authors say minimizes the expected squared approximation error of the layer output. The last frozen layer is kept intact so the shape of the representation entering the trainable top layers does not change. The full procedure is given as Algorithm 1, with TOA as Algorithm 2.
The convergence part formalizes the intuition: three assumptions (smooth objective, bounded gradient variance across clients, bounded error from freezing) are used to bound the drop in the global loss per round and to characterize how close the algorithm gets to a stationary point as a function of the learning rate.
Why This Matters
Impact on research: The paper challenges the assumption that any layer-freezing pattern delivers memory savings. It shows that the direction of freezing determines whether memory is actually saved, which reframes layer freezing as a memory-management technique rather than only a compute/communication technique. It also introduces TOA to federated learning as an alternative to quantization for compressing transmitted parameters.
Real-world applications (as framed by the paper and its setting):
- IoT deployments where many heterogeneous devices must jointly train a model without centralizing data.
- Face recognition services running on edge devices.
- Audio analysis applications on edge hardware.
- Any edge deployment where device memory, rather than compute or bandwidth, is the binding constraint — devices that would otherwise be excluded from federated training entirely.
Industry relevance: Hardware fleets in IoT, mobile, and embedded settings are heterogeneous; a scheme that assigns each device a number of frozen layers according to its capacity lets low-capability devices participate instead of being dropped (the paper notes that excluded devices cause "severe information loss"). Reduced upload/download volume also lowers bandwidth costs for operators, and avoiding full-model storage on device lowers the memory bill for constrained hardware.
Future Directions
Note: the paper states that Section VI concludes the work and lists future directions, but that section's content is not included in the available text, so the authors' own proposed next steps are not reported here. Questions the work naturally raises include:
- How should the number of frozen layers
l_kbe chosen automatically from a device's measured memory, rather than assumed as given input to the algorithm? - How should the TOA scaling factor
sbe tuned to balance accuracy against communication, and how does it interact with the convergence bound's dependence on the freezing-error termD? - Can the ordered-freezing idea be combined with client sampling, partial participation, or the dropout-based methods the paper criticizes, and how do those combinations behave under non-iid data?
- How does FedOLF behave under more severe forms of device heterogeneity or on architectures other than CNNs, such as transformers, where layer redundancy patterns may differ?
Target Audience
Researchers and practitioners working on federated learning, edge/IoT machine learning, and on-device training under resource constraints. It is most useful for readers already comfortable with neural network training internals (forward and backward propagation, activations, gradients) and with basic federated learning terminology (clients, rounds, aggregation, non-iid data). Engineers optimizing models for memory-limited embedded hardware will find the practical argument clearer than the theory; readers interested in convergence guarantees for non-convex federated optimization will find the assumptions and theorems the substantive part.
Authors’ abstract
Federated Learning (FL) has emerged as a privacy-preserving paradigm for training machine learning models across distributed edge devices in the Internet of Things (IoT). By keeping data local and coordinating model training through a central server, FL effectively addresses privacy concerns and reduces communication overhead. However, the limited computational power, memory, and bandwidth of IoT edge devices pose significant challenges to the efficiency and scalability of FL, especially when training deep neural networks. Various FL frameworks have been proposed to reduce computation and communication overheads through dropout or layer freezing. However, these approaches often sacrifice accuracy or neglect memory constraints. To this end, in this work, we introduce Federated Learning with Ordered Layer Freezing (FedOLF). FedOLF consistently freezes layers in a predefined order before training, significantly mitigating computation and memory requirements. To further reduce communication and energy costs, we incorporate Tensor Operation Approximation (TOA), a lightweight alternative to conventional quantization that better preserves model accuracy. Experimental results demonstrate that over non-iid data, FedOLF achieves at least 0.3%, 6.4%, 5.81%, 4.4%, 6.27% and 1.29% higher accuracy than existing works respectively on EMNIST (with CNN), CIFAR-10 (with AlexNet), CIFAR-100 (with ResNet20 and ResNet44), and CINIC-10 (with ResNet20 and ResNet44), along with higher energy efficiency and lower memory footprint.