Research
DepthBench: Measuring How Residual Connections Enable More Computational Depth
Overview Research area: Transformer architecture design for large language models, specifically how the choice of residual connection and normalization scheme determines whether added depth produces u

- arXiv
- 2609.32534
- Published
- 2026-09-26
- Authors
- Keyu Wang, Yangyi Huang, Jiale Kang, David González-Martínez, Weiyang Liu, Shiwei Liu
AI summary
Overview
- Research area: Transformer architecture design for large language models, specifically how the choice of residual connection and normalization scheme determines whether added depth produces useful computation.
- Technical level: Advanced. The paper assumes familiarity with Pre-LN Transformers, residual streams, RMSNorm, learning-rate sweeps, and layer-level interpretability diagnostics.
- Scope in one sentence: DepthBench is a controlled, iso-parameter benchmark that sweeps the width–depth aspect ratio (d_model / n_layer) across 10 residual/normalization designs to test which ones turn architectural depth into effective computational depth.
What This Paper Is About
Adding layers is an obvious way to give a Transformer more computational capacity, but in standard Pre-LN models the later layers often contribute little, a problem the paper calls the "curse of depth" or "Pre-LN dilution." Recent fixes (LayerNorm Scaling, AttnRes, HC, mHC) claim better depth utilization, but they are usually compared using different architectures, training recipes, and system setups, so it is unclear whether the gains really come from depth being used more effectively. DepthBench isolates this by holding parameter count and the pre-training recipe fixed while trading width for depth, then measuring whether each architecture's extra layers actually do useful work.
Key Contributions
-
A controlled iso-parameter benchmark for depth. DepthBench varies the aspect ratio d_model / n_layer from shallow–wide to deep–narrow while approximately matching the parameter budget and keeping the pre-training recipe, tokenizer, and optimizer fixed, across 10 representative architectures.
-
Four complementary experimental suites. A main 400M benchmark (7 shapes, aspect ratios 9.1–76.0), an iso-backbone control at roughly 300M backbone parameters (6 shapes, aspect ratios 6.5–78.0) to remove the confound that narrower models spend a larger fraction of the budget on the Transformer backbone rather than embeddings and LM head, a multi-scale suite at {200M, 300M, 400M, 500M} parameters with three shapes (≈28, ≈43, ≈76), and a validation-at-scale suite at approximately 1.6B parameters with aspect ratios 27.9, 43.2, and 73.1.
-
Architecture ranking with per-configuration learning-rate tuning. Every aspect ratio is swept over several learning rates and reported at its best-performing rate, so differences reflect architecture rather than a suboptimal learning rate (optimal rates for the 400M models are reported in Table 4).
-
Layer-level diagnostics of depth utilization. The paper measures angular distance between representations at different depths, pairwise causal scores and permutation scores, explicit residual-path weights, single-layer pruning on ARC-Easy, and LogitLens KL divergence, to explain why some architectures benefit from depth.
Main Findings
-
Standard residual designs do not benefit from depth under a fixed budget. For Pre-LN at 400M parameters, validation loss monotonically increases from 2.759 at L = 16 to 2.782 at L = 32, meaning shifting capacity from width to depth hurts. Pre-LN's optimum sits at a relatively large aspect ratio; the paper reports optima in the range 42.7–76.0 (for example d_model = 1120 with n_layer = 20).
-
Normalization- and scaling-based variants are weak or non-monotonic. Sandwich-LN, LNS, DeepNorm, KEEL, and MoDA show either weak trends or non-monotonic ones, with their optima at intermediate or shallower shapes rather than at deep, narrow shapes.
-
HC and Full AttnRes keep improving as models get deeper and narrower. Full AttnRes improves from 2.751 at L = 16 to 2.718 at L = 32, and HC improves from 2.729 at L = 16 to 2.699 at L = 32 — the opposite trend from Pre-LN. The improvement continues out to an extreme aspect ratio of 9.1 with d_model = 640 and n_layer = 70.
-
Depth gains are not merely a backbone-size effect. With the backbone size fixed at about 300M parameters, deep–narrow HC and Full AttnRes models improve even as total model size decreases, indicating the benefit comes from the width–depth allocation itself.
-
Derived variants inherit less of the benefit. Block AttnRes is competitive but peaks at an intermediate shape, reaching 2.715 at L = 20 before degrading to 2.730 at L = 32. mHC achieves strong overall performance with its best result at L = 24 but shows no consistent gain with increasing depth.
-
Gains transfer to domain-specific evaluation. Teacher-forced negative log-likelihood on coding (MBPP and HumanEval), STEM (SciQ and GPQA), and math (GSM8K and MATH-500) is generally lower for deeper HC and Full AttnRes models, with the clearest trends on STEM and math; Pre-LN and most other variants show weaker or non-monotonic trends.
-
The trend persists at larger scale for Full AttnRes, but not cleanly for HC. Across 200M–500M models Full AttnRes consistently benefits from deeper–narrower shapes and this persists at 1.6B. HC is less stable: the 500M run at the largest aspect ratio encounters gradient explosion, and the 1.6B results do not show the same clear improvement with depth, which the paper attributes to HC's sensitivity to optimization hyperparameters.
-
Representation diversity separates the architectures. Using angular distance between hidden states at layers ℓ and ℓ+n, Pre-LN, Sandwich-LN, LNS, DeepNorm, and KEEL produce representations that become broadly similar across depth. HC and AttnRes show larger angular distances and a markedly non-smooth structure, with changes that stay layer-specific and non-uniform. Block AttnRes shows a block-structured pattern (smooth within blocks, sharp changes across block boundaries), and mHC shows a moderate non-smooth pattern.
-
Late layers are functionally more consequential in HC and Full AttnRes. Restricting analysis to the last three quarters of the network, for the L = 32 models only about 1% of Pre-LN's late-layer causal scores exceed 0.45, versus roughly 16% for Full AttnRes and 7% for HC. Under layer permutation, about 8% of Pre-LN scores exceed 0.15, versus approximately 46% for Full AttnRes and 35% for HC.
-
The two winning architectures use depth differently. Pre-LN accumulates all updates in a single residual stream with weight w ≡ 1. Full AttnRes forms a normalized, non-negative mixture over the embedding and all preceding sublayer outputs, giving direct access to earlier computations, while HC realizes cross-depth access recursively, with the contribution of branch output f_i to branch ℓ given by w = β_iᵀ A_{i+1} ⋯ A_{ℓ-1} α_ℓ. Full AttnRes's learned paths are strongest locally per the described behavior for HC; Full AttnRes directly retrieves stored earlier outputs whereas HC propagates and recombines them.
-
Layer pruning and LogitLens show different depth profiles. On ARC-Easy accuracy, Pre-LN and HC are both dominated by the first layer and lose only modestly when most later layers are pruned, whereas Full AttnRes is sensitive across a broader range including both its first and final layers. LogitLens KL divergence to the final prediction decreases monotonically for Pre-LN, follows a similarly smooth trajectory for HC (with earlier representations farther from the final prediction), and is markedly less monotonic for Full AttnRes, consistent with continued retrieval and late integration of stored sub-layer outputs.
-
Block AttnRes has an identifiable design cost. With the number of boundaries fixed at 8 and block size growing with depth (b = 4, 5, 6, 7, 8 for L = 16, 20, 24, 28, 32), the depth-mixing softmax initially places essentially zero weight on the running sum, so layers effectively re-read the embedding. This dead segment scales with depth at 2, 5, and 8 layers for L = 16, 24, and 32, meaning a quarter of the L = 32 network cannot benefit from stacking layers.
-
mHC's residual transport is more constrained. When all attention and FFN residual maps are composed along the forward pass, the resulting products have consistently lower effective rank in mHC, 1.44–1.65, compared with 2.53 (the excerpt cuts off the label for the comparison value). These statistics are averaged over sampled held-out FineWeb-Edu data, with 95% document-bootstrap intervals shown as bands and error bars.
-
Deep models carry a systems-level cost. Although deeper architectures can improve modeling performance, they increase compute and memory overhead while reducing hardware utilization — a trade-off the paper states requires better infrastructure and kernel optimization to resolve at scale.
Methodology in Plain English
The core move is to treat depth as a capacity allocation decision rather than as free extra compute. The authors fix an approximate parameter budget and then choose, for each target number of layers L, a hidden dimension d so that the model hits that budget. Since total parameters scale roughly as 12·L·d² + 2·V·d (V being vocabulary size), going deeper necessarily means going narrower, producing a clean spectrum from shallow–wide to deep–narrow shapes.
All ten architectures are instantiated on a shared LLaMA-like backbone so that only the normalization or residual design differs. Sub-1B models use multi-head self-attention with 16 heads, RoPE, RMSNorm with ε = 10⁻⁶, SwiGLU feed-forward layers, the GPT-NeoX tokenizer with a vocabulary of 50280, and untied input/output embeddings. The 1.6B models follow the Qwen3-1.7B attention configuration with grouped-query attention (16 query heads, 8 key-value heads). For sub-1B models the intermediate dimension is ceil₁₆(8d/3); for 1.6B models it is 3d.
Pre-training uses OLMo-core from scratch on FineWeb-Edu with a token budget of 20 tokens per parameter following Chinchilla (8B tokens for 400M models, 32B for 1.6B models), sequence length 2048, and a global batch size of 512 sequences (about 1M tokens per step). Optimization uses AdamW with β₁ = 0.9, β₂ = 0.95, ε = 10⁻⁸, weight decay 0.1, gradient clipping at 1.0, and OLMo-core default initialization with standard deviation 0.02. The learning rate follows cosine decay with 10% linear warmup down to 10% of peak. For 400M models the learning-rate sweep is {5×10⁻⁴, 1×10⁻³, 2×10⁻³, 5×10⁻³}, except LNS which uses {2×10⁻³, 5×10⁻³, 1×10⁻², 2×10⁻²}; these are reused for the iso-backbone and 200M–500M experiments, and 5×10⁻⁴ is used for the 1.6B models. The reported optimal learning rates are 2×10⁻³ for every architecture except LNS (1×10⁻²) and DeepNorm (1×10⁻³).
For the mechanistic analysis, the authors measure cosine-based angular distance between representations at different depths, causal scores (how much skipping a layer changes a later layer's update) and permutation scores (how much performance drops when two layers are swapped), both aggregated only over the last three quarters of the network with thresholds of 0.45 for causal scores and 0.15 for permutation scores. They also trace explicit residual-path weights, prune individual layers and measure ARC-Easy accuracy, and apply LogitLens to compare intermediate predictions to the final one. For the angular-distance analysis they use a unified learning rate of 2×10⁻³ for all models (only LNS departs from its optimal rate) because the learning rate substantially affects learned representations, and they exclude DeepNorm because it explodes at 2×10⁻³.
Why This Matters
The paper reframes depth scaling from "add more layers" to "choose an architecture whose extra layers actually compute something," and it provides evidence that the residual-connection design, not just the normalization tweak, decides which regime a model lands in. For research, it supplies a controlled protocol that separates depth effects from capacity and recipe confounds, and it argues that reported gains of new residual schemes should be checked for whether they reflect effective computational depth or confounded factors.
Real-world implications:
- Model-family design. Groups choosing how to split a fixed training budget between width and depth now have evidence that Pre-LN's optimum lies at wide shapes while HC and Full AttnRes favor deep–narrow ones.
- Pre-training infrastructure. Because deeper models raise compute and memory overhead and lower hardware utilization, teams would need better kernels and systems work before deep-shape gains translate into practical training runs.
- Domain-specific model selection. Since lower NLL on coding, STEM, and math tracks the depth-friendly architectures, teams targeting those domains have a concrete criterion for evaluating residual designs.
- Diagnostics for existing checkpoints. The causal, permutation, pruning, and LogitLens diagnostics offer a way to check whether later layers in an already-trained model are doing meaningful work.
Industry relevance is direct: the paper notes that related residual designs are being adopted in frontier LLMs such as Kimi K3 and DeepSeek V4, and it overlays its findings with the model sizes and aspect ratios of representative industry-released Pre-LN models, using shaded regions to illustrate the contrasting shape preferences.
Future Directions
- Better infrastructure for deep models. The paper identifies a systems-level efficiency trade-off and states that realizing deep-shape benefits at scale requires better infrastructure and kernel optimization.
- Stability and hyperparameter sensitivity of HC. HC encounters a gradient explosion at 500M and loses its clear depth trend at 1.6B, which the authors attribute to sensitivity to optimization hyperparameters; finer per-scale learning-rate tuning and the stability-oriented constraints used by mHC are natural follow-ups.
- Fixing the Block AttnRes schedule. The initial dead segment (2, 5, and 8 layers at L = 16, 24, and 32) means a quarter of the L = 32 network cannot benefit from stacking layers, raising the question of whether a different boundary or block-size schedule would recover depth scaling.
- Whether the trends extend past 1.6B. The validation-at
Authors’ abstract
Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (\text{e.g.}, LayerNorm Scaling) or residual connections (\text{e.g.}, mHC, AttnRes) to enable better information flow and depth utilization. However, it remains unclear whether they truly translate increased architectural depth into effective computational depth, and whether their reported gains stem from better access to information across depth, or unaccounted-for confounding factors. In this paper, we introduce \textbf{DepthBench}, a controlled benchmark for studying computational depth across various architectures. We systematically vary the width--depth aspect ratio ($d_{\text{model}}/n_{\text{layer}}$) from shallow--wide to deep--narrow shapes, while keeping the model size and pre-training recipe fixed. Across 10 representative architectures, we find that the benefit of allocating more capacity to depth is strongly architecture-dependent. Standard Pre-LN and most of its norm- and scaling-based variants provide little benefit and can even degrade performance as models become deeper and narrower, whereas HC and Full AttnRes improve consistently even at extreme deep shapes. These gains extend beyond pre-training loss and consistently translate into improved domain-specific performance and effective computation. Controlled layer-level analyses further show that the gains of HC and Full AttnRes are associated with more effective utilization of additional layers, revealing distinct mechanisms of computational depth across architectures. Overall, our results identify residual connection design as a key determinant of whether depth can serve as a meaningful scaling axis by enabling additional architectural depth to translate into effective computation.