Picture for Yan-Lun Chen

Yan-Lun Chen

RAS: Measuring LLM Safety Through Refusal Alignment

Add code
Jun 24, 2026
Viaarxiv icon

Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution

Add code
Jun 24, 2026
Viaarxiv icon

IU: Imperceptible Universal Backdoor Attack

Add code
Feb 28, 2026
Viaarxiv icon

Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge

Add code
Feb 27, 2025
Figure 1 for Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
Figure 2 for Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
Figure 3 for Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
Figure 4 for Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
Viaarxiv icon

BADTV: Unveiling Backdoor Threats in Third-Party Task Vectors

Add code
Jan 04, 2025
Viaarxiv icon