Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

Nov 16, 2025

Haotian Jin, Yang Li, Haihui Fan, Lin Shen, Xiangfang Li, Bo Li

Figure 1 for Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

Figure 2 for Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

Figure 3 for Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

Figure 4 for Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

Share this with someone who'll enjoy it:

Abstract:Backdoor attacks pose a serious threat to the security of large language models (LLMs), causing them to exhibit anomalous behavior under specific trigger conditions. The design of backdoor triggers has evolved from fixed triggers to dynamic or implicit triggers. This increased flexibility in trigger design makes it challenging for defenders to identify their specific forms accurately. Most existing backdoor defense methods are limited to specific types of triggers or rely on an additional clean model for support. To address this issue, we propose a backdoor detection method based on attention similarity, enabling backdoor detection without prior knowledge of the trigger. Our study reveals that models subjected to backdoor attacks exhibit unusually high similarity among attention heads when exposed to triggers. Based on this observation, we propose an attention safety alignment approach combined with head-wise fine-tuning to rectify potentially contaminated attention heads, thereby effectively mitigating the impact of backdoor attacks. Extensive experimental results demonstrate that our method significantly reduces the success rate of backdoor attacks while preserving the model's performance on downstream tasks.

View paper on

Share this with someone who'll enjoy it:

Title:Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

Paper and Code