Papers
arxiv:2610.00348

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

Published on Sep 29
· Submitted by
George Drayson
on Oct 2
Authors:
,
,

Abstract

Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.

Community

Paper author Paper submitter

NEEDLE is a training-free backdoor defence designed to remove a backdoor with minimal changes to model behaviour and safety. It estimates a backdoor direction (how the trigger shifts the model's activations) and a refusal subspace (directions that mediate refusal), then edits the model's weights layer by layer: each layer's attention and MLP output weights are orthogonalised against the backdoor direction while keeping their refusal projections fixed, and a closed-form correction keeps the activations' refusal projections unchanged as earlier layers are edited.

Codebase: https://github.com/LocaiLabs/NEEDLE/
Models: https://huggingface.co/collections/locailabs/needle

This is great work, super relevant in this day and age!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.00348
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 24

Browse 24 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.00348 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.00348 in a Space README.md to link it from this page.

Collections including this paper 1