Skip to content
AI.info

Research

Scaling depth capacity via zero/one-layer model expansion

Scaling Depth Capacity via Zero/One-Layer Model Expansion Overview Research area: Efficient large-scale deep learning training — specifically progressive training (also called model expansion or model

Scaling depth capacity via zero/one-layer model expansion
arXiv
2511.04981
Published
2025-11-07
Authors
Zhiqi Bu

AI summary

Scaling Depth Capacity via Zero/One-Layer Model Expansion

Overview

Research area: Efficient large-scale deep learning training — specifically progressive training (also called model expansion or model growth), combining optimization theory, feature learning theory (muP), and large language model pre-training.

Technical level: Advanced. The paper derives a convergence bound under convex, G-Lipschitz loss, uses spectral scaling conditions from muP theory to justify hyperparameter transfer, and assumes familiarity with warmup-stable-decay (WSD) learning rate schedules and transformer architectures.

One-sentence scope: The paper proposes expanding a model's depth from a zero-layer or one-layer source model into a much deeper target model, and shows this saves roughly 80% of pre-training compute while matching the loss of full-size training.

What This Paper Is About

Training a large model costs compute proportional to its size (approximately 6BTN FLOPs), so training a small model first and growing it later should be cheaper — but it is unclear how to initialize new layers, how to set the learning rate, and when to grow. This paper argues that growing depth from a model with zero or one transformer layer is the most compute-efficient option, and it supplies theory and recipes for making that growth work. It also identifies a "mixing behavior" in which the progressive run's loss catches up to a fully trained fixed-size run, which is what allows aggressive late expansion.

Key Contributions

  1. Analysis of depth expansion as an initialization problem that preserves feature learning. The author shows that this framing enables hyperparameter transfer (e.g. the learning rate) across the whole progressive run, rather than requiring new hyperparameter tuning after expansion.

  2. Identification of the learning rate schedule as a critical ingredient. The warmup-stable-decay (WSD) schedule is shown both theoretically and empirically to improve convergence for progressive training, in contrast to cosine decay.

  3. Discovery of mixing behaviors. Progressive training's loss and training dynamics mix with those of fixed-size training, which supports transferring the mixing time across settings and using single-stage expansion rather than the multi-stage expansion used in prior work.

  4. Demonstration that zero/one-layer progressive training gives the best tradeoff between computation and loss. Scaling laws are validated across GPT2, LLAMA3, DeepSeekV3 and others, with the compute-efficiency advantage growing at larger scales.

  5. A convergence theory of progressive training under convex optimization, yielding insights on initialization, learning rate schedule, and a projected-gradient-descent (PGD) interpretation.

Main Findings

  • Compute savings on GPT2: Zero/one-layer progressive training on GPT2 saves approximately 80% compute, equivalently an approximately 5× acceleration, while attaining a loss comparable to a fully trained 60-layer model with 7B parameters.

  • Loss gap is small: In Figure 1, the difference in final validation loss versus fixed-size training is less than 0.5% for the 124M runs and less than 0.2% for the 7B runs. For full runs, depth expansion happens at 80% of iterations; for early-stopped runs, expansion happens at 2% of iterations, where warmup ends.

  • Scaling law gains: On LLAMA3 (dense, token-per-param = 50) and DeepSeekV3 (MoE, token-per-param = 100) pre-training on OpenWebText under WSD, progressive training shows a 3 to 5× improvement in compute efficiency, with an increasing advantage at larger scales. Scaling laws for those models use sizes from 0.25B to 2B for LLAMA3 and 0.2B to 0.5B active parameters for DeepSeekV3, with depth expansion at 80% of iterations.

  • Random and copying are the best initializations: Across GPT2 (dense, MHA), LLAMA3 (dense, GQA), Qwen3 (dense, GQA), DeepSeekV3 (MoE, MLA) and Mixtral (MoE, GQA), random and copying initialization of new layers work best. Pure zero initialization is invalid because it kills gradient flow and makes new layers untrainable.

  • Function-preserving conflicts with feature learning: Table 1 summarizes the tradeoff — copying and random are not function-preserving but have high trainability and enable feature learning; zero initialization is function-preserving but has low trainability and blocks feature learning. Copying and random satisfy the muP spectral scaling condition, while zero and copying_zero (with a zero sub-layer) do not.

  • WSD beats cosine for late expansion: For ResNet, the small model can still catch up with the large model when expansion happens at τ ≈ 0.8T under WSD, but fails around τ ≥ 0.7T under cosine. For GPT, the small model catches up until τ ≈ 0.8T under WSD but fails around τ ≥ 0.5T under cosine.

  • Mixing time is robust to τ under WSD: Expanding a GPT 1-layer model at 10% or 60% of the horizon both need approximately 16B tokens, or 30k iterations, to mix with 12-layer training under WSD. Expanding at 80% of the horizon cannot mix well because the learning rate has decayed. The paper's own WSD setup uses 2% warmup, 10% decay and 528k iterations at constant learning rate; subtracting approximately 40k iterations of mixing time places depth expansion at t = 480k.

  • Single-stage expansion suffices: Across 150 runs (3 large model sizes, 5 small model sizes, 10 expansion times) targeting {12, 24, 36}-layer GPT2 with {124M, 400M, 1B} parameters, zero/one-layer progressive training nearly captures the Pareto-optimal loss-compute tradeoff. Expanding from 1-layer or from 6-layer at τ/T ≈ 0.6 is similarly effective, but the latter is far more expensive. Multi-stage expansion (e.g. 0 → 2 → 12) brings no efficiency or loss benefit because it decomposes into two single-stage expansions.

  • Ordering matters only for multi-layer expansion: For copying-based multi-layer expansion, copying all layers is consistently better than copying only the last layer; copying_inter and copying_stack are almost indistinguishable. These orderings are irrelevant for zero-layer expansion and equivalent for one-layer expansion.

  • Progressive training is more than a point on the tradeoff curve: Under an identical compute budget (a grown model trained for 120k iterations after expansion at τ = 0.8T, compared against a fixed-size run of the same 120k iterations with the same schedule), progressive training converges much faster, i.e. it genuinely inherits progress from the small model rather than merely moving along the loss-compute frontier.

  • Appending new random layers at the bottom works best: Inserting randomly initialized layers at the bottom yields much smaller loss spikes than inserting at the top. Random initialization works well on GPT2 and MoE but slightly less well on ResNet, which the author attributes to intermittent insertion across ResNet's four inhomogeneous stages.

  • copying_zero variants differ: copying_zeroN (zeroed normalization sub-layers) has weak trainability, whereas copying_zeroL (zeroed last linear sub-layer) converges as fast as plain copying and avoids any loss spike.

  • MoE depth expansion works: The recipe transfers to MoE models such as DeepSeekV3, and the paper distinguishes this from MoE upcycling, which increases width/sparsity but not depth and which has reported negative results because the grown MoE becomes worse than a from-scratch MoE after a few hundred billion tokens.

  • Theoretical characterization: Under convex, G-Lipschitz loss, the gap between the progressive-training bound and the fixed-size bound reduces to two terms: a weighted minimum-loss term and a term involving the distance of the extra parameters from their optimum. Minimizing the weighted term favors smaller learning rates before expansion than after, which is consistent with WSD rather than with learning rate decay.

Methodology in Plain English

The author takes a residual network and splits its "depth" into the repeated blocks while keeping the embedding, LM head, or ResNet stage machinery. A zero-layer model is just [Embedding, LM_head]; a one-layer GPT2 has N = 1; the ResNet analogies are ResNet14 [1,1,1,1] for zero-layer and ResNet26 [2,2,2,2] for one-layer. The small model is trained, then new layers are inserted (by random initialization or by copying existing layers), and training continues with the same hyperparameters.

The study compares several ways of creating the new layers (copying, random, zero, copying_zero) and several orderings when many layers must be inserted (copying all layers, copying only the last layer, stacking, interpolating). It then varies the learning rate schedule (cosine versus WSD), the timing of expansion, the size of the source model, and the number of expansion stages. Models are trained on OpenWebText at 1024 sequence length using the nanoGPT codebase, and ResNets are trained on ImageNet at 224×224 for 100 epochs, with Muon-NSGD as the main optimizer at 0.01 weight decay and no gradient clipping.

On the theory side, the author writes down the standard descent inequality for SGD on the small model and the large model, telescopes the two phases around the expansion step τ, and compares the resulting bound against the pure fixed-size bound. To make the comparison tractable, the large model is written as a concatenation of the small model and extra parameters, W = [w, x]. This yields an interpretation of progressive training as projected gradient descent that masks the deeper layers to zero, followed by a "teleportation" of the extra parameters from zero to a good initialization, then ordinary SGD.

Why This Matters

Impact on research. The paper argues that the field has largely evaluated progressive training from the perspective of the grown model versus the target model, which hides the mixing behavior and overstates speedups by ignoring the cost of the small model. Re-plotting from that older perspective removes the observed mixing. It also shows that function-preserving initialization — long a goal of depth expansion work — is in tension with feature learning, and that the muP-based hyperparameter transfer resolves the practical tuning problem that otherwise makes expansion expensive.

Real-world applications.

  • Pre-training large language models where the compute budget is the binding constraint, since the source model can be as small as zero or one layer.
  • Training sparse mixture-of-experts models, which the paper validates on DeepSeekV3 and Mixtral configurations.
  • Vision model training, validated on ResNet (ImageNet, 224×224, 100 epochs), where zero-layer equates to ResNet14 and one-layer to ResNet26.
  • Reducing the environmental and cost footprint of large training runs, motivated in the paper by LLAMA-4's over 7M GPU hours and an estimated 2,000 tons of carbon emissions.

Industry relevance. The recipe is deliberately practical: keep the same optimizer (Muon-NSGD or another muP-scaled optimizer) and the same hyperparameters before and after expansion, use WSD, and determine the expansion time from two small early-stopped runs. That avoids re-tuning learning rates at each scale, which is often the hidden cost of growth-based training.

Future Directions

  • Scaling width and depth together. The conclusion explicitly suggests pushing the efficiency frontier by expanding both width and depth, not depth alone.
  • Extending the theory beyond convex, G-Lipschitz losses. The convergence analysis assumes convexity and Lipschitz continuity; the paper motivates this by noting that deep learning training dynamics resemble convex optimization, but a non-convex analysis remains open.
  • Determining the mixing time more cheaply. The current prescription requires two small-scale early-stopped runs to locate the mixing time; a predictive rule would reduce that overhead.
  • Clarifying copying_inter versus copying_stack for multi-layer expansion. The paper notes prior work disagrees on which is better and that it avoids the question entirely by expanding from zero or one layer, leaving the question unresolved for deeper source models.

Target Audience

Researchers and engineers working on large-scale model pre-training, efficient training, and training-dynamics theory. It is most useful to practitioners who can afford to restructure a training pipeline (choosing optimizers, schedules, and expansion timing) and to theorists interested in convergence bounds and feature-learning conditions for progressive training. Readers without a background in muP-style scaling or convex optimization bounds will find Sections 3.2 and 4 dense, but the recipe in Section 7 is self-contained enough for applied use.

Authors’ abstract

Model depth is a double-edged sword in deep learning: deeper models achieve higher accuracy but require higher computational cost. To efficiently train models at scale, progressive training (also known as model expansion) scales up model capacity during training and significantly reduces computation with little performance degradation. In this work, we study the depth expansion of large-scale models through the lens of optimization theory and feature learning, offering insights on the initialization of new layers, hyperparameter transfer, learning rate schedule, and timing of model expansion. Specifically, we propose zero/one-layer progressive training to achieve an optimal tradeoff between computation and loss, with a comprehensive ablations on our expansion strategy. For example, zero/one-layer progressive training on GPT2 can save $\approx 80\%$ compute, or equivalently achieve an $\approx 5\times$ acceleration, while attaining a loss comparable to that of a fully trained 60-layer model with 7B parameters, thus demonstrating a mixing behavior in terms of loss. Furthermore, scaling laws on LLAMA3 and DeepSeekV3 models show a $3\sim 5\times$ improvement in compute efficiency, with an increasing advantage at larger scales.

Read the original paper