Research
Compressing LLMs with MoP: Mixture of Pruners
Overview Research area: Efficient machine learning — specifically structured pruning of Large Language Models (LLMs) and vision-language models. Technical level: Intermediate. The paper assumes famili

- arXiv
- 2602.06127
- Published
- 2026-02-05
- Authors
- Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias, Leandro Giusti Mugnaini, Keith Ando Ogawa, Lucas Pellicer, Rosimeire Pereira Costa, Edson Bollis, Anna Helena Reali Costa, Artur Jordao
AI summary
Overview
Research area: Efficient machine learning — specifically structured pruning of Large Language Models (LLMs) and vision-language models.
Technical level: Intermediate. The paper assumes familiarity with transformer architecture (attention heads, MLP layers, projection matrices) and standard compression terminology, but the core idea is explained without requiring deep theory.
Scope: The paper proposes MoP (Mixture of Pruners), an iterative pruning framework that combines depth pruning (removing whole transformer layers) and width pruning (removing attention heads and MLP neurons) into a single compression path, and evaluates it on LLaMA 7B, LLaMA-2 7B, LLaMA-3 8B and the multimodal LLaVA-1.5 7B.
What This Paper Is About
Structured pruning methods for LLMs generally pick one axis: they either delete entire layers (depth pruning), which gives strong inference speedups but coarse control, or delete internal components such as attention heads and MLP neurons (width pruning), which allows fine-grained selection but limited acceleration. The paper argues this dichotomy forces a trade-off and instead asks whether a method can alternate between both dimensions during a single pruning run. MoP is the proposed answer: at each iteration it builds a depth candidate and a width candidate using the same parameter budget, scores them, and keeps whichever one better preserves the original model, repeating until the target compression is reached.
Key Contributions
- MoP, a hybrid pruning framework. An iterative method that unifies depth and width pruning. At every iteration it generates two branches — one layer removed versus an equivalent number of parameters removed in width — and a path criterion selects which candidate advances.
- A modular design. MoP accepts plug-in choices for the layer criterion, the width criterion, and the path criterion, so it can absorb future state-of-the-art pruning criteria without redesign.
- A fair parameter-matched comparison. Because layer removal produces coarse jumps in model size, MoP recomputes at each iteration the fraction of parameters represented by the selected layer and prunes that same fraction in width, so both candidates are evaluated on a consistent basis.
- Extension to multimodal models and a new empirical observation. MoP is applied to LLaVA-1.5 7B, and the authors report being the first to observe that text-only recovery fine-tuning can restore performance on visual benchmarks in a pruned vision-language model.
Main Findings
- MoP leads at every tested compression tier on both language models. In Table 2, MoP ranks first at 20%, 30% and 40% compression on LLaMA-2 7B (mean averages 65.22 ± 0.21, 61.54 ± 0.27 and 55.37 ± 0.45) and on LLaMA-3 8B (65.44 ± 1.04, 60.43 ± 0.59 and 55.34 ± 1.42). The dense baselines are 68.99 for LLaMA-2 7B and 72.70 for LLaMA-3 8B.
- Robustness across random paths. MoP's worst-seed mean on LLaMA-2 7B is 64.97% at 20%, 61.24% at 30% and 54.97% at 40%. These exceed the best reported baseline means at each tier — AmoebaLLM at 20% (64.90%) and AMP at 30% and 40% (61.02% and 53.06%). The worst-seed advantage grows from 0.07 percentage points at 20% to 1.91 pp at 40%.
- The margin widens on LLaMA-3 8B. Worst-seed means are 64.33% at 20%, 59.87% at 30% and 54.35% at 40%, exceeding the strongest baselines by 1.40 pp at 20% (CoMe at 62.93%), 1.86 pp at 30% (AMP at 58.01%) and 1.70 pp at 40% (AMP at 52.65%). Against the third-place method per tier, the margin expands from 1.58 pp at 20% (LINEARPATCH) to 5.53 pp at 30% and 6.55 pp at 40% (both Yang et al.).
- Mixing beats each dimension alone. When MoP is restricted to prune only in depth or only in width, the mixed strategy consistently outperforms both single-dimension variants across all compression ratios on LLaMA-2 7B, which the authors describe as a Pareto improvement over single-dimension approaches (Figure 2).
- Random path selection works as well as metric-based selection. At 30% compression on LLaMA 7B, the average accuracy was 59.21 for KL, 60.68 for cosine similarity, 60.84 for perplexity and 60.83 ± 0.43 for random selection (mean and standard deviation over three runs). The authors adopt random selection, treating it as a conservative lower bound, and report results over three independent random pruning paths thereafter.
- Real latency reduction. On a single NVIDIA RTX 4090 GPU, with a prompt of 12 input tokens generating 128 output tokens at batch size 1, the dense LLaMA-2 7B baseline latency is 2.21s. Pruned models reach 1.82 ± 0.01s, 1.60 ± 0.03s and 1.36 ± 0.03s at 20%, 30% and 40% compression, i.e. speedups of 1.22 ± 0.01×, 1.38 ± 0.02× and 1.63 ± 0.04×, and end-to-end latency reductions of 18.0 ± 0.7%, 27.5 ± 1.1% and 38.7 ± 1.5% respectively. The abstract states a 39% latency reduction at 40% compression. At 20% compression MoP outperforms AMP (1.19×) and Shortened LLaMA (1.13×), and its 1.38× at 30% exceeds the 1.17× Yang et al. report at the same tier and their 1.30× at 50% compression.
- Multimodal performance is largely preserved. On LLaVA-1.5 7B (Table 3), the dense model scores 68.02 on ScienceQA, 53.52 on VizWiz, 28.44 on MM-Vet and 59.80 on LLaVA-Bench, for a mean of 52.44. At 20% compression the mean is 46.70 ± 0.96, at 30% it is 44.61 ± 1.98 and at 40% it is 40.86 ± 1.22. At 20% compression the model retains 81.79% of its original capability on MM-Vet and 91.97% on LLaVA-Bench; at 40% it preserves 77.92% of its mean predictive capabilities.
- Text-only recovery fine-tuning restores visual-task performance. Over the first 17 MoP iterations on LLaVA-1.5, the gap between fine-tuned and non-fine-tuned models peaks at 23.71 pp, and while unrefined models show a standard deviation peak of 8.85 pp, fine-tuned variants limit it to 2.80 pp.
- Comparable degradation across modalities. At 30% compression, LLaVA-1.5's mean performance drop is 7.83 pp relative to its dense baseline, closely paralleling the 7.45 pp drop observed in LLaMA-2 7B at the same rate.
Methodology in Plain English
The starting point is a transformer decoder: a stack of layers, each containing an attention module and an MLP, with all layers sharing the same dimensions. That uniformity means a whole layer is the smallest depth unit that can be removed without breaking the architecture, while inside a layer the attention heads and MLP neurons can be removed individually for much finer control, because their internal dimensions run into the thousands.
MoP runs a loop until the target compression ratio is met. Each pass, it does the following. It first picks a layer to delete using a simple rule derived from prior work: the deeper layers of a transformer tend to be the most redundant, with the exception of the very last ones, so MoP always removes the third-to-last layer, keeping the final two intact and working from the end toward the beginning. It then measures how much of the model's total parameter count that layer represents, and uses that same fraction as the width pruning ratio for the second candidate. For width pruning it uses AMP, which scores attention heads and MLP neurons by the magnitude of their activations and prunes uniformly across layers.
Both candidates — the layer-pruned one and the width-pruned one — get a short recovery fine-tuning pass, and are then compared against the original unpruned model by a path criterion. That criterion outputs a single scalar measuring how far the candidate has drifted from the original, with lower meaning better. The authors test cosine similarity between output logits, KL divergence between output token distributions, perplexity on calibration data, and a random selector. Crucially, whichever candidate wins, the algorithm advances with the pre-fine-tuning version of it; the short fine-tuning exists only to make the comparison meaningful. The loop repeats, and once the target compression is reached, the surviving model gets a full recovery fine-tuning pass.
Practical details: calibration and metric evaluation use the WikiText-2 test set, with 128 random non-empty texts for cosine similarity and KL divergence, and the full test split in fixed 2048-token segments for perplexity. Because capability degrades progressively during pruning, the authors apply a dynamic alignment rule — at pruning iteration i, they run 10 × i fine-tuning steps on Alpaca. Final recovery fine-tuning uses Alpaca via LoRA, with rank 32 and alpha 10 for language models, learning rate 3e-4, batch size 16 and two epochs with AdamW; for multimodal models only the rank and alpha change, to 8 and 16. Training runs on an RTX 4090, where one MoP run to a target compression ratio, including recovery fine-tuning, takes 2 hours. Evaluation uses EleutherAI LM Harness for the language benchmarks (standard accuracy for WinoGrande, normalized accuracy for the rest) and LMMs-Eval for the multimodal ones.
Why This Matters
Impact on research. The paper challenges an implicit assumption in structured pruning: that a method must commit to either depth or width. By showing that randomized alternation between the two already beats either one alone, it suggests that the value comes from distributing removal across dimensions rather than from sophisticated scoring. Its claim that a worst-seed MoP path beats the best reported baseline mean also reframes how pruning results should be read — seed selection stops being an implicit hyperparameter. The finding that text-only recovery fine-tuning repairs visual-task performance in a pruned vision-language model opens a question that the paper explicitly notes is under-explored relative to LLM pruning.
Real-world applications:
- Deploying capable LLMs on single-GPU or edge hardware with tight latency budgets, where a 1.38× speedup at 30% compression is directly useful.
- Real-time and interactive assistants that need to generate long responses (128 output tokens in the paper's latency setup) within strict per-request time limits.
- Multimodal assistants built on vision-language models such as LLaVA, where the method preserves visual reasoning scores at 20-40% compression.
- Reducing the energy cost of large-scale model serving, which the authors frame as a Green AI contribution alongside the latency benefit.
Industry relevance. The method is modular, so an organization already using a particular width or depth pruning criterion can drop MoP on top without replacing its existing tooling. The reported 2-hour full run, including recovery fine-tuning, on a single RTX 4090 lowers the practical barrier to reproducing and adopting the approach, and the authors release code and models publicly.
Future Directions
- Stronger path criteria as drop-in replacements. The paper deliberately adopts random path selection as a "conservative lower bound" and states that MoP is modular enough to support stronger future criteria. Which criteria would meaningfully exceed random selection remains open.
- Extending to other multimodal and larger architectures. Only LLaVA-1.5 7B was tested in the multimodal setting, and language evaluation covers models up to LLaMA-3 8B. Whether the multimodal findings generalize is not established.
- Understanding why text-only fine-tuning repairs visual tasks. The authors report the phenomenon and call the multimodal pruning domain under-explored, but the paper does not explain the mechanism behind the cross-modal recovery.
- Generalizing the speedup evidence. Latency was measured only on a single NVIDIA RTX 4090 with one prompt/generation configuration and batch size 1; behavior on other hardware, batch sizes and serving stacks is not reported.
Target Audience
This paper is most useful to machine learning engineers and researchers working on model compression, efficient inference, or LLM deployment, particularly those who need to shrink models to fit specific latency or memory budgets. It is also relevant to practitioners working with vision-language models who want to understand how structured pruning behaves across modalities, and to researchers interested in pruning criteria and evaluation methodology, given the paper's argument for worst-seed reporting. Readers without prior exposure to transformer internals or pruning terminology would benefit from background reading first.
Authors’ abstract
The high computational demands of Large Language Models (LLMs) motivate methods that reduce parameter count and accelerate inference. In response, model pruning emerges as an effective strategy, yet current methods typically focus on a single dimension-depth or width. We introduce MoP (Mixture of Pruners), an iterative framework that unifies these dimensions. At each iteration, MoP generates two branches-pruning in depth versus pruning in width-and selects a candidate to advance the path. On LLaMA-2 and LLaMA-3, MoP advances the frontier of structured pruning, exceeding the accuracy of competing methods across a broad set of compression regimes. It also consistently outperforms depth-only and width-only pruning. Furthermore, MoP translates structural pruning into real speedup, reducing end-to-end latency by 39% at 40% compression. Finally, extending MoP to the vision-language model LLaVA-1.5, we notably improve computational efficiency and demonstrate that text-only recovery fine-tuning can restore performance even on visual tasks.