Skip to content
AI.info

Research

Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning

Overview Research area: Natural Language Processing, specifically structured pruning of Transformer-based models (attention head pruning) for efficient inference and deployment. Technical level: Inter

Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
arXiv
2510.13832
Published
2025-10-10
Authors
Minsik Choi, Hyegang Son, Changhoon Kim, Young Geun Kim

AI summary

Overview

Research area: Natural Language Processing, specifically structured pruning of Transformer-based models (attention head pruning) for efficient inference and deployment.

Technical level: Intermediate. The core idea is intuitive (combine two existing scores), but the paper supports it with a risk-decomposition analysis featuring lemmas on loss-increase bounds, generalization gaps, and gradient orthogonality.

Scope: The paper introduces HIES (Head Importance–Entropy Score), a pruning criterion that combines gradient-based Head Importance Score (HIS) with Attention Entropy (AE), and evaluates it on BERT_base, LLaMA-2 7B, ViT_Large, and LLaVA-1.5 7B across GLUE, commonsense reasoning, image classification, and vision-language benchmarks.

What This Paper Is About

Large Transformer models carry many layers and attention heads, which makes inference expensive. Existing head-pruning methods rank heads by HIS, a gradient-based measure of each head's contribution to the loss, but HIS captures only the magnitude of gradient-driven contribution and ignores how a head distributes attention across tokens. The paper's goal is a single pruning score that combines gradient importance with attention-distribution structure, so that pruning stays accurate and stable even at aggressive compression ratios.

Key Contributions

  1. A unified pruning criterion (HIES). The paper defines HIES as a weighted combination of normalized HIS and normalized inverse attention entropy — HIES_h = α·HIŜ_h + (1−α)(1 − AÊ_h), with α ∈ [0,1) — and applies min–max normalization to both metrics before combining them.

  2. A theoretical risk decomposition justifying the combination. The analysis separates pruning risk into a loss-increase term controlled by HIS (Lemma 4.1, an upper bound that becomes ∑(1−m_h)·HIS_h + ⅛‖W^O‖²₂ ∑(1−m_h)‖A_h‖²_F under binary cross-entropy) and a generalization-gap term upper-bounded via attention entropy deficit AD_h(x) = 1 − AE_h(x) ∈ [0,1].

  3. Optimality and complementarity results. Lemma 4.2 shows that selecting the k heads with the largest HIES values is the globally optimal solution to the cardinality-constrained pruning objective; Lemma 4.3 shows the gradients of HIS and AE are orthogonal in expectation, framing them as complementary (magnitude vs. dispersion) rather than redundant signals.

  4. Empirical validation across models and modalities. HIES is tested against Random, L2-Norm, HIS, Attention Deficit (AD), LLM-Pruner (channel-wise and block-wise), and SliceGPT, on language, vision, and vision-language tasks.

Main Findings

  • Average quality gain on GLUE with BERT_base: HIES improves model quality by 8.21% on average relative to the reported baselines (Table 1).

  • Stability gain: Table 2 reports a 3.3% average stability gain over HIS, measured as stability against the unpruned model.

  • Aggressive-pruning gains: At a 50% pruning ratio, the paper reports gains of up to 15.2% in model quality and 2.04× in stability; Section 5.2 describes these as relative to the best-performing baseline, while the abstract frames them as gains over HIS-only methods.

  • Low vs. high sparsity: At ratios of 10% or below, HIS and HIES perform comparably and HIS is marginally better in a few cases; at ratios of 30% and above, HIES is more stable and yields flatter accuracy–sparsity curves.

  • Distinct head-removal patterns: HIS prunes mainly from lower layers, producing an approximately bottom-up pattern, while HIES selects more dispersedly across lower, middle, and upper layers (Figure 3, CoLA, 30%–50% pruning).

  • Scalability to LLaMA-2 7B: Across HellaSwag, Winogrande, ARC-e, ARC-c, and OBQA at 10–60% pruning, HIES improves accuracy by up to +10.54% and stability by up to +6.21% relative to HIS, averaged over tasks. Gains are noted to persist across all pruning ratios, including on ARC-c, the hardest benchmark in that suite.

  • Vision and vision-language transfer: On ViT_Large at 20% sparsity, HIES reaches 82.70% average accuracy versus 55.20% for HIS, described as 49.82% higher accuracy; at 50% sparsity on VizWiz-VQA and MM-Vet, HIES improves by 11.71% over HIS. At 10% sparsity HIES is slightly below HIS (−0.56% average on ViT_Large, −2.84% on LLaVA-1.5 7B).

  • Motivating diagnosis: In a BERT analysis on SST-2, some heads pruned for low HIS still assigned high attention to sentiment-discriminative tokens, while some retained high-HIS heads allocated strong attention to non-informative tokens — evidence that HIS alone misses token-level attention structure.

  • Results deferred to appendices: GSM8K and MMLU results are stated to be in Appendix D.6, and CIFAR-100, Food-101, and Fashion-MNIST results in Appendix D.7; those numbers are not reported in the main text. The paper states that code will be released upon publication.

Methodology in Plain English

The authors start from an existing pruning signal, HIS, which is computed by measuring how much the loss would change if a given attention head were switched off, using gradients. They observe that this score can be nearly identical for two heads with very different behavior — one that focuses sharply on a few informative tokens and one that spreads attention broadly. To capture that difference, they compute attention entropy for each head, which measures how evenly the head spreads its attention across input tokens.

They normalize both scores to the range 0 to 1 with min–max scaling, because the two quantities live on different scales and are not directly comparable. Then they combine them into a single score with a tunable weight α: a head scores high if it is both gradient-important and low-entropy (sharply focused). The paper tunes α per task and per compression setting, and analyzes α sensitivity in Appendix D.8. Heads are then ranked by HIES and the lowest-scoring ones are removed.

To justify the combination, the authors split the risk of pruning into a loss-increase term — bounded by HIS plus a quadratic correction involving the output projection norm and head activation norms — and a generalization-gap term that depends on the attention entropy deficit. They prove that selecting the k highest-HIES heads minimizes the resulting composite risk bound, and that the gradient directions of HIS and AE are orthogonal in expectation, so the two signals carry non-overlapping information.

Experiments use publicly available fine-tuned BERT_base checkpoints, a LLaMA-2 7B checkpoint from Hugging Face, ViT_Large, and LLaVA-1.5 7B, with HIES computed using a calibration size of 32. Quality is measured by task metrics (accuracy, Matthews correlation, F1, Pearson correlation), and stability is measured relative to the unpruned model.

Why This Matters

Impact on research. The paper reframes head pruning as a two-signal problem — magnitude of loss contribution plus structure of attention — and grounds the combination in a risk bound and an orthogonality result, rather than proposing an empirical heuristic. It also positions attention entropy, previously used to detect entropy collapse in training, as an inference-time stability indicator.

Real-world applications:

  • Real-time translation on mobile and edge devices, where latency and memory budgets are tight.
  • Intelligent voice assistants that must run under constrained compute.
  • Personalized recommendation systems deployed on consumer-grade hardware.
  • Vision-language assistants and multimodal question answering, where the paper's LLaVA-1.5 7B results on VizWiz-VQA and MM-Vet are directly relevant.

Industry relevance. Head pruning preserves layer topology while cutting attention FLOPs and KV-cache memory, which the paper notes simplifies checkpoint compatibility and serving integration. This makes HIES-style criteria easier to slot into existing LLM compression pipelines than methods that alter model structure more broadly. The stability claim matters commercially: models that degrade unpredictably under input or compression shifts are harder to deploy in production.

Future Directions

  • Extending beyond head pruning. The conclusion proposes extending the approach toward broader structured sparsity, which the paper does not test.
  • Reducing the α tuning burden. The trade-off hyperparameter is currently tuned per task and per compression setting; whether a single α generalizes across settings is an open question, with sensitivity analysis deferred to Appendix D.8.
  • Reasoning-heavy and larger-scale evaluation. The main text evaluates classification-style and multimodal tasks, with GSM8K and MMLU results deferred to Appendix D.6 — broader reasoning coverage and larger models remain to be explored.
  • Combining with complementary compression techniques. The paper notes its results hold across attention variants, but does not report combinations with quantization or other compression methods that are typically stacked with pruning in deployment pipelines.

Target Audience

Researchers and engineers working on model compression, efficient inference, and deployment of large language and vision-language models. It is most useful to readers already familiar with Transformer attention mechanics and structured pruning terminology, and least suited to complete beginners, given the lemma-driven theoretical analysis. Practitioners building compression pipelines for resource-constrained or edge deployment are the primary beneficiaries.

Authors’ abstract

Transformer-based models have achieved remarkable performance in NLP tasks. However, their structural characteristics-multiple layers and attention heads-introduce efficiency challenges in inference and deployment. To address these challenges, various pruning methods have recently been proposed. Notably, gradient-based methods using Head Importance Scores (HIS) have gained traction for interpretability, efficiency, and ability to identify redundant heads. However, HIS alone has limitations as it captures only the gradient-driven contribution, overlooking the diversity of attention patterns. To overcome these limitations, we introduce a novel pruning criterion, HIES (Head Importance-Entropy Score), which integrates head importance scores with attention entropy, providing complementary evidence on per-head contribution. Empirically, HIES-based pruning yields up to 15.2% improvement in model quality and 2.04x improvement in stability over HIS-only methods, enabling substantial model compression without sacrificing either accuracy or stability. Code will be released upon publication.

Read the original paper