Research
DART-ing Through the Drift: Dynamic Tracing of Knowledge Neurons for Adaptive Inference-Time Pruning
Overview Research area: Natural Language Processing, efficient LLM inference, and structured pruning of transformer Feed-Forward Networks (FFNs). Technical level: Advanced. Scope: The paper introduces

- arXiv
- 2601.22632
- Published
- 2026-01-30
- Authors
- Abhishek Tyagi, Yunuo Cen, Shrey Dhorajiya, Bharadwaj Veeravalli, Xuanyao Fong
AI summary
Overview
Research area: Natural Language Processing, efficient LLM inference, and structured pruning of transformer Feed-Forward Networks (FFNs). Technical level: Advanced. Scope: The paper introduces DART (Dynamic Attention-Guided Runtime Tracing), a training-free runtime pruning method that adapts FFN neuron masks as autoregressive context evolves, and evaluates it on LLaMA-3.2-3B, LLaMA-3.1-8B, and additional models in the appendix.
What This Paper Is About
Large Language Models contain substantial parameter redundancy, especially in FFNs, but existing pruning methods depend on dataset-specific calibration and apply static masks that cannot adapt during generation. The paper identifies a failure mode called knowledge drift, where neurons inactive during prefill later become critical for factual or domain-specific content, and proposes DART to dynamically update neuron-level masks by monitoring shifts in attention score distributions.
Key Contributions
- A lightweight, model-agnostic pruning method that derives per-layer masks for FFN sublayers by modeling each layer’s geometric contribution to the residual stream, achieving deterministic inference speedups while preserving accuracy.
- Identification of knowledge drift, a failure mode in dynamic pruning where static masks derived from an initial input prefix fail to support the evolving semantic requirements of autoregressive generation, leading to error accumulation during long-horizon tasks.
- An online knowledge drift detector that monitors distributional shifts in attention outputs and acts as a runtime supervisor, triggering adaptive mask updates to resolve semantic misalignment.
- Extensive benchmarks across zero-shot and multi-shot domain-specific datasets and multi-topic summarization, showing DART outperforms static, dynamic, structured, and unstructured pruning baselines.
Main Findings
- Accuracy gains at 70% FFN sparsity: Across ten benchmarks, DART outperforms the prior dynamic baseline DejaVu. The introduction reports up to +14.5 accuracy gains on LLaMA-3.2-3B and up to +19.6 on LLaMA-3.1-8B at 70% FFN sparsity; the abstract separately reports accuracy gains of up to 14.5% on LLaMA-3.1-8B at 70% FFN sparsity.
- Zero-shot examples on LLaMA-3.2-3B at 70% FFN sparsity: BoolQ Dense 74.04, Wanda 66.27, DejaVu 44.25, DART 53.24; HellaSwag Dense 74.13, Wanda 38.23, DejaVu 26.41, DART 52.77; ARC-e Dense 72.05, Wanda 45.50, DejaVu 25.71, DART 50.38.
- Zero-shot examples on LLaMA-3.1-8B at 70% FFN sparsity: BoolQ Dense 83.09, Wanda 67.79, DejaVu 42.14, DART 66.20; HellaSwag Dense 79.32, Wanda 44.99, DejaVu 26.68, DART 64.58; ARC-e Dense 82.53, Wanda 50.34, DejaVu 25.08, DART 59.43.
- Domain-specific multi-shot tasks: Five-shot accuracy remains competitive, though the gap with Wanda can shrink to approximately 5%. On LLaMA-3.2-3B MMLU, Dense 58.14, Wanda 26.60, DejaVu 23.94, DART 28.51. On LLaMA-3.1-8B MMLU, Dense 66.61, Wanda 29.70, DejaVu 25.20, DART 34.14.
- Summarization: DART achieves up to 3× better ROUGE-L scores with respect to static-masked pruning on summarization tasks, with performance comparable to the original dense models.
- Generation with knowledge tracing: Using 500 prompts with average prompt length 35 tokens and max generation length 500 tokens, sparse without tracing versus sparse plus tracing yields ROUGE-L 0.27 → 0.38, BLEU 0.20 → 0.36, BERTScore (F1)
Authors’ abstract
Large Language Models (LLMs) exhibit substantial parameter redundancy, particularly in Feed-Forward Networks (FFNs). Existing pruning methods suffer from two primary limitations. First, reliance on dataset-specific calibration introduces significant data dependency and computational overhead. Second, being predominantly static, they fail to account for the evolving subset of knowledge neurons in LLMs during autoregressive generation as the context evolves. To address this, we introduce DART, i.e., Dynamic Attention-Guided Runtime Tracing), a lightweight, training-free method that performs on-the-fly context-based pruning. DART monitors shifts in attention score distributions to infer context changes, dynamically updating neuron-level masks to retain salient parameters. Across ten benchmarks, DART outperforms prior dynamic baseline, achieving accuracy gains of up to 14.5% on LLAMA-3.1-8B at 70% FFN sparsity. Furthermore, DART achieves up to 3x better ROUGE-L scores with respect to static-masked pruning on summarization tasks, with its performance comparable to the original dense models. We conclusively demonstrate that the proposed framework effectively adapts to diverse semantic contexts, preserves model capabilities across both general and domain-specific tasks while running at less than 10MBs of memory for LLAMA-3.1-8B(16GBs) with 0.1% FLOPs overhead. The code is available at https://github.com/seeder-research/DART.