Skip to content
AI.info

Research

Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction

Overview Research area: Computer vision, specifically Vision Transformer (ViT) architecture design and parameter efficiency. Technical level: Intermediate. The paper assumes familiarity with transform

Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction
arXiv
2512.01059
Published
2025-11-30
Authors
Anantha Padmanaban Krishna Kumar

AI summary

Overview

Research area: Computer vision, specifically Vision Transformer (ViT) architecture design and parameter efficiency.

Technical level: Intermediate. The paper assumes familiarity with transformer blocks, MLP sublayers, and standard ImageNet training protocols, but its methods are simple enough to follow without deep theoretical background.

Scope: A controlled empirical comparison of two MLP parameter-reduction strategies (weight sharing versus width reduction) applied to ViT-B/16 trained on ImageNet-1K, evaluating accuracy, training stability, and efficiency trade-offs.

What This Paper Is About

Scaling laws and prior results generally suggest that larger Vision Transformers perform better, which encourages the assumption that adding parameters is the main path to accuracy. This paper tests the opposite direction: it removes 32.7% of the parameters from the MLP blocks of ViT-B/16 and checks whether ImageNet-1K accuracy and training stability hold up or improve. The goal is to determine whether standard ViT-B/16 sits in an overparameterized regime where MLP capacity can be cut without harm, and to compare two mechanistically different ways of doing that cutting at matched parameter counts.

Key Contributions

  1. Demonstrating that, for ViT-B/16 on ImageNet-1K, two parameter-reduction schemes in the MLPs improve accuracy and training stability despite removing 32.7% of parameters.
  2. Comparing parameter sharing (GroupedMLP) against width reduction (ShallowMLP) at matched parameter counts to clarify their practical trade-offs in compute, memory, and throughput.
  3. Connecting these findings to the broader literature on overparameterization in transformers and outlining open questions about architectural constraints in ViTs.
  4. Releasing all code at https://github.com/AnanthaPadmanaban-KrishnaKumar/parameter-efficient-vit-mlps

Main Findings

  • Both reduced models beat the baseline: GroupedMLP reaches 81.47% top-1 accuracy and ShallowMLP reaches 81.25%, versus 81.05% for the 86.6M-parameter ViT-B/16 baseline. GroupedMLP improves by 0.42% (p=0.018) and ShallowMLP by 0.20% (p=0.031). Both variants use 58.2M parameters, or 67.3% of the baseline count.

  • Top-5 accuracy tracks the same pattern: Baseline 95.36% ± 0.08, GroupedMLP 95.66% ± 0.08, ShallowMLP 95.52% ± 0.02.

  • Training stability improves sharply: The peak-to-final accuracy gap falls from 0.47% ± 0.04 for the baseline to 0.06% ± 0.06 for GroupedMLP and 0.03% ± 0.01 for ShallowMLP. Across seeds, the baseline degrades by 0.43–0.50%, while GroupedMLP stays within 0.12% and ShallowMLP within 0.04% of peak.

  • The two designs offer different trade-offs: GroupedMLP keeps the baseline computational cost at 16.9 GFLOPs and a 4× expansion ratio while reducing memory footprint by 4.3%. ShallowMLP cuts compute to 11.3 GFLOPs with a 2× expansion ratio, achieving 38% higher throughput (1,411 img/s versus 1,020 for the baseline) and 29% lower memory usage. GroupedMLP measures 1,017 img/s.

  • Optimization follows a different trajectory: The baseline reaches 80% accuracy by epoch 185 and peaks at epoch 219 ± 13, then degrades continuously with rising loss. GroupedMLP and ShallowMLP peak roughly 50 epochs later (272 ± 1 and 273 ± 0), maintaining stable accuracy and loss through training completion.

  • Parameter efficiency improves: Both strategies achieve approximately 49% higher accuracy per parameter than the baseline (1.40 and 1.39 versus 0.94 accuracy per million parameters). ShallowMLP reaches 7.20 accuracy per GFLOP versus 4.80 for the baseline, a 50% improvement.

  • Two distinct mechanisms produce similar gains: Weight sharing (reducing 12 unique MLPs to 6) and width reduction (halving the hidden dimension from 3072 to 1536) both inherit initialization from the full 86.6M-parameter model, which the author suggests may matter more than the eventual parameter count.

Methodology in Plain English

The author starts from the official timm implementation of ViT-B/16 and applies architectural changes after initialization, so all models begin from identical weights. Two variants are built, each dropping 32.7% of parameters:

GroupedMLP makes adjacent transformer blocks share the same MLP weights. Blocks (2i, 2i+1) for i in {0,...,5} reference identical parameters, collapsing 12 unique MLPs into 6. To keep gradient flow healthy, shared parameters are scaled by 1/√2 at initialization (applied to W_fc1, W_fc2, and b_fc1), which preserves forward-pass variance. Attention layers remain independent.

ShallowMLP instead halves the MLP hidden dimension from 3072 to 1536 in every block while keeping parameters independent. Rather than initializing a narrow MLP from scratch, the author slices the corresponding weight matrices out of the full ViT-B/16 — W_fc1 takes the first d/2 rows, W_fc2 takes the first d/2 columns — preserving the initialization statistics of the larger model.

Both models are trained on ImageNet-1K for 300 epochs with a standard recipe: AdamW (β1=0.9, β2=0.999, weight decay 0.05), cosine learning rate starting at 10⁻³ with 5-epoch warmup, batch size 1024, and augmentations including MixUp 0.8, CutMix 1.0, RandAugment, and DropPath 0.1. An exponential moving average of weights with decay 0.9998 is used for evaluation. Experiments run over seeds {42, 123}, with significance tested via paired t-tests. Evaluation covers validation accuracy at the best checkpoint, peak-to-final accuracy gap as a stability metric, inference throughput and memory, and accuracy per parameter and per FLOP.

Why This Matters

Impact on research: The result challenges the assumption that parameter count is a reliable proxy for effective capacity in Vision Transformers. It provides concrete evidence that architectural constraints — parameter sharing and reduced width — can function as useful inductive biases rather than pure compression techniques, and it connects to existing work on deep double descent, sparse subnetworks, ALBERT-style sharing, and documented redundancy in ViT attention and patch representations.

Real-world applications:

  • Memory-constrained deployment, where GroupedMLP's 4.3% memory reduction at unchanged 16.9 GFLOPs helps fit models onto limited hardware.
  • Latency-sensitive inference, where ShallowMLP's 38% throughput gain and 29% lower memory suit real-time vision systems.
  • Cloud-scale serving, where a 50% improvement in accuracy per GFLOP reduces compute cost per inference without losing accuracy.
  • Fine-tuning pipelines built on ViT-B/16 checkpoints, where a smaller, more stable MLP sublayer could lower training cost and reduce degradation near the end of long schedules.

Industry relevance: Vision Transformers are widely deployed as backbones, and the paper offers two simple, drop-in post-initialization modifications requiring no iterative pruning. The observed stability improvement — an order-of-magnitude reduction in peak-to-final degradation — matters for practitioners who currently lose accuracy late in training and may benefit from more predictable convergence.

Future Directions

  1. Scaling studies across model sizes: Test whether the gains hold from ViT-Tiny to ViT-Huge, and how the optimal degree of parameter reduction depends on total capacity.
  2. Broader architecture and modality coverage: Extend the study to DeiT, Swin, convolutional networks, hybrid models, and language transformers, as well as fine-tuning and transfer-learning setups.
  3. Richer sharing schemes: Explore sharing patterns beyond adjacent pairs, different group sizes, and alternative scaling rules for shared weights to find more effective or flexible designs.
  4. Loss landscape and gradient analysis: Investigate whether the stability gains correspond to flatter minima, and how information flow and gradient dynamics differ between shared and narrowed MLPs. The paper notes that detailed ablations isolating the effect of large-model initialization statistics are left to future work.

Target Audience

Researchers and practitioners working on Vision Transformer efficiency, model compression, and architecture design. It is most useful to engineers deploying ViT-B/16 under memory or latency constraints, and to researchers interested in overparameterization and training stability in transformers. Readers seeking a rigorous theoretical account of why reduced MLP capacity helps will find empirical observations and hypotheses here rather than proofs — the paper explicitly notes that confirming mechanisms such as flatter minima will require more direct theoretical and empirical analysis.

Authors’ abstract

Although scaling laws and many empirical results suggest that increasing the size of Vision Transformers often improves performance, model accuracy and training behavior are not always monotonically increasing with scale. Focusing on ViT-B/16 trained on ImageNet-1K, we study two simple parameter-reduction strategies applied to the MLP blocks, each removing 32.7\% of the baseline parameters. Our \emph{GroupedMLP} variant shares MLP weights between adjacent transformer blocks and achieves 81.47\% top-1 accuracy while maintaining the baseline computational cost. Our \emph{ShallowMLP} variant halves the MLP hidden dimension and reaches 81.25\% top-1 accuracy with a 38\% increase in inference throughput. Both models outperform the 86.6M-parameter baseline (81.05\%) and exhibit substantially improved training stability, reducing peak-to-final accuracy degradation from 0.47\% to the range 0.03\% to 0.06\%. These results suggest that, for ViT-B/16 on ImageNet-1K with a standard training recipe, the model operates in an overparameterized regime in which MLP capacity can be reduced without harming performance and can even slightly improve it. More broadly, our findings suggest that architectural constraints such as parameter sharing and reduced width may act as useful inductive biases, and highlight the importance of how parameters are allocated when designing Vision Transformers. All code is available at: https://github.com/AnanthaPadmanaban-KrishnaKumar/parameter-efficient-vit-mlps.

Read the original paper