Skip to content
AI.info

Research

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

Overview Research area: LLM security and safety (arXiv category cs.CR) — specifically backdoor/data-poisoning attacks on large language models and the defence methods that try to remove them. Technica

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
arXiv
2610.00348
Published
2026-09-29
Authors
Minoo Kim, Vasileios Lampos, George Drayson

AI summary

Overview

  • Research area: LLM security and safety (arXiv category cs.CR) — specifically backdoor/data-poisoning attacks on large language models and the defence methods that try to remove them.
  • Technical level: Advanced. The paper assumes familiarity with residual-stream activations, activation steering / steering vectors, singular value decomposition, ridge regression, LoRA adapters, and fine-tuning-based defences.
  • Scope: The paper proposes Needle, a training-free backdoor removal method that permanently edits model weights via sequential orthogonalisation against an estimated backdoor direction while preserving a rank-4 refusal subspace, evaluated on Gemma-3-4B-IT and Qwen3-4B-Instruct-2507 (with additional experiments on Gemma-3-12B-IT) across sentiment steering, targeted refusal, and code injection attacks.

Authors: Minoo Kim, Vasileios Lampos (Centre for AI, Computer Science, UCL), George Drayson (Locai Labs). The preprint is noted as currently under peer review.

What This Paper Is About

Backdoor attacks implant a hidden trigger in an LLM during training so that the model behaves normally on ordinary prompts but produces an attacker-chosen response — hostile language, targeted refusal, or malicious code — when the trigger appears. Existing defences either require extra fine-tuning or intervene at inference time with a clean reference model, and they can shift the model's output distribution, degrading both performance and safety. This paper's goal is to remove the backdoor through a one-off, permanent edit to the model weights that suppresses the trigger behaviour while explicitly leaving refusal behaviour intact.

Key Contributions

  1. Empirical characterisation of backdoors as a direction. The authors demonstrate that backdoor behaviour can largely be captured by a single direction in activation space, and that this direction is highly correlated with directions that mediate refusal — meaning naïve removal damages safety.
  2. The Needle method. A novel backdoor removal method that orthogonalises model weights against the backdoor direction while retaining the safety (refusal behaviour) of the original model, using no clean reference model and no access to the original poisoned training data.
  3. Four-perspective evaluation. The method is evaluated against existing work on removal effectiveness, distribution shift, and the resulting change in model performance and safety, across multiple model families and attack types.
  4. State-of-the-art removal with minimal side effects. Needle achieves the lowest mean Attack Success Rate among the evaluated defences, including 0% on challenging code injection attacks, while producing the lowest KL divergence and minimal changes in capability and safety.

Main Findings

  • Backdoor directions overlap refusal directions. Figure 1 shows that removing the backdoor direction alone, or while preserving only a single refusal direction (k=1), raises harmfulness on the BadNet sentiment steering attack; much of the overlap between the backdoor direction and the leading refusal directions lies outside the k=1 direction, which motivates preserving a rank-4 refusal subspace (k=4).
  • Lowest mean Attack Success Rate. Needle reduces mean ASR from 99.50% (± 0.71) to 1.67% (± 2.41) on Gemma-3-4B-IT and from 99.08% (± 0.84) to 5.00% (± 6.16) on Qwen3-4B-Instruct-2507, the lowest among evaluated defences.
  • Baselines are unreliable at removal. Mean ASR for the baselines ranges between 27.75% and 69.25%. The one exception is CROW on Gemma, which removes the backdoor (5.00% ASR) but loses on average 14.80% of relative capability and increases harmful responses by 19.78 points.
  • Minimal capability loss. Needle's average relative capability loss is 0.48% (± 0.74) on Gemma and 0.72% (± 0.90) on Qwen, versus 14.80% (± 4.10) for CROW on Gemma and 14.80%-scale losses elsewhere.
  • Smallest distribution shift. KL divergence from the backdoored to the defended model is 0.03 (± 0.02) for Needle on Gemma and 0.11 (± 0.09) on Qwen, compared with values such as 0.54 (± 0.17) for SFT and 0.64 (± 0.34) for CROW on Gemma, and 0.57 (± 0.22) / 0.69 (± 0.29) for SFT / BD-VAX on Qwen. The paper states the fine-tuning process induces larger shifts in the model's output distribution than Needle.
  • 0% ASR on code injection. On the domain-specific code injection attack (inserting a secret API key in triggered code-generation prompts), Needle reaches 0% ASR, as claimed in the abstract and contributions.
  • ATR stays low. Needle's Accidental Trigger Rate (target responses on untriggered prompts) is 1.00 (± 1.15) on Gemma and 1.17 (± 1.07) on Qwen, comparable to the baselines, which range from 0.25 to 0.75 on Gemma and 0.25 to 0.75 on Qwen.
  • Triggered and untriggered activations converge. Figure 3 reports that after Needle, triggered and untriggered responses become indistinguishable when projected onto the backdoor direction and refusal subspace from layer 30 onwards, while harmful responses maintain their refusal component.

Methodology in Plain English

  1. Estimate a backdoor direction. For each layer, the authors collect mean activations on triggered prompts and on ordinary prompts and take their difference. They subtract the component along the ordinary-prompt mean (so the direction is orthogonal to it) and normalise. The trigger is assumed to have already been identified, so the defender can construct triggered prompts by querying the model.
  2. Build a refusal subspace rather than a single refusal direction. Using mean activations from refused and compliant responses to harmful prompts, they compute a difference vector, remove its projection onto the midpoint of those two means, and normalise. They then form several refused/compliant pairs, centre each difference, strip out the components along the midpoint and the mean refusal direction, concatenate the results, and take an SVD. Three leading right singular vectors plus the mean refusal direction give an orthonormal rank-4 refusal basis per layer.
  3. Edit the weights, not the activations. The whole point is a permanent change to the model, so no inference-time steering or reference model is needed. At each layer they compute the part of the backdoor direction that lies outside the refusal subspace and subtract the projection of the weight matrix along that component — a closed-form update that removes the backdoor projection while leaving the refusal projection unchanged.
  4. Correct the drift layer by layer. Because editing earlier layers changes later activations, the weight-level guarantee is not enough for activation-level behaviour. They orthogonalise layers 12–34 in increasing order, recomputing activations after each edit, and fit a second, ridge-regularised update to the MLP output matrix that restores the refusal projections while remaining constrained to be orthogonal to the backdoor direction.
  5. Test broadly. Backdoored models are built by SFT with a LoRA adapter (and separately with full parameter fine-tuning) on clean plus poisoned data, across two behaviours from the BackdoorLLM benchmark (sentiment steering with the response "You are stupid!" and targeted refusal with "I'm sorry, but as an AI assistant, I do not have the capability to follow the given instruction"), two trigger types (BadNet and Sleeper), and one custom code injection attack. Needle is compared against SFT, OSFT, CROW, and BD-VAX, and measured on ASR, ATR, capability change (MMLU, GSM8K, HellaSwag, ARC-Challenge, IFEval, HumanEval, MBPP), safety change (WildGuardMix test-set harmful prompts), and KL divergence on clean prompts.

Why This Matters

Backdoors in open-weight LLMs are hard to find and hard to remove: they persist through later safety training, can be implanted with relatively few poisoned examples, and the models can be redistributed as checkpoints without training data. Most existing defences change the model's behaviour on ordinary prompts as a side effect, which is a safety problem in itself. Needle's relevance is that it removes the trigger behaviour while explicitly protecting refusal, and its low KL divergence suggests the rest of the model is largely undisturbed.

Real-world applications:

  • Vetting third-party open-weight checkpoints. A defender who downloads a redistributed model but has no access to its training data can, given a known trigger, strip the backdoor with a weight edit rather than retraining.
  • Code assistants. The paper's code injection attack — an inserted secret API key in triggered code-generation prompts — is a direct model of supply-chain leakage in AI coding tools.
  • Safety-critical deployments. Because Needle preserves the refusal subspace, a de-backdoored model can keep refusing harmful requests, and its harmful-response rate change is small and in one case negative (−3.74 ± 11.24 on Qwen).
  • Compliance and audit workflows. Weight-level removal is a permanent artefact that can be verified and shipped, unlike inference-time wrappers that add computation and require a surrogate model.

Industry relevance: the method targets the open-weight distribution ecosystem, where model provenance is uncertain, and it addresses the compute constraint that inference-time defences impose — a practical concern for anyone serving LLMs at scale. The paper provides code and models as linked artefacts.

Future Directions

  • Combining with trigger detection. Needle assumes the trigger and its insertion rule are already known, so the paper positions it as complementary to methods that detect or recover poisoned triggers; tighter integration with detection is a natural next step.
  • Beyond the studied attack families. The paper focuses on data poisoning, while the attack taxonomy it cites also includes weight poisoning and hidden-state or chain-of-thought manipulation. Extending the weight-orthogonalisation idea to weight-level and activation-level attacks is open.
  • Scaling and architecture transfer. Additional experiments with Gemma-3-12B-IT validate generalisability at larger parameter counts, but the full picture beyond the tested sizes and model families is not established in the reported content.
  • Adaptive attackers and subspace size. The choice of a rank-4 refusal subspace is shown to matter (k=1 leaves harmfulness), which raises the question of how the rank should be chosen and whether an attacker who anticipates this removal could evade it.

Target Audience

Researchers and engineers working on LLM security, red-teaming, and model editing; practitioners who distribute or deploy open-weight models and need to certify that a checkpoint is free of implanted behaviour; and readers already comfortable with activation steering and weight-space interventions, since the method's derivation and evaluation assume that background. Those new to backdoor attacks will find the problem framing accessible but the methodology itself demanding.

Authors’ abstract

Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.

Read the original paper