Picture for Xue Tan

Xue Tan

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Add code
Aug 10, 2026
Viaarxiv icon

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

Add code
Jul 07, 2026
Viaarxiv icon

Knowledge Database or Poison Base? Detecting RAG Poisoning Attack through LLM Activations

Add code
Nov 28, 2024
Figure 1 for Knowledge Database or Poison Base? Detecting RAG Poisoning Attack through LLM Activations
Figure 2 for Knowledge Database or Poison Base? Detecting RAG Poisoning Attack through LLM Activations
Figure 3 for Knowledge Database or Poison Base? Detecting RAG Poisoning Attack through LLM Activations
Figure 4 for Knowledge Database or Poison Base? Detecting RAG Poisoning Attack through LLM Activations
Viaarxiv icon