Picture for Xuying Li

Xuying Li

Steering Vision-Language Models with Joint Sparse Autoencoders

Add code
Jun 24, 2026
Viaarxiv icon

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

Add code
Jun 09, 2026
Viaarxiv icon

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

Add code
Aug 14, 2025
Viaarxiv icon

Targeting the Core: A Simple and Effective Method to Attack RAG-based Agents via Direct LLM Manipulation

Add code
Dec 05, 2024
Figure 1 for Targeting the Core: A Simple and Effective Method to Attack RAG-based Agents via Direct LLM Manipulation
Figure 2 for Targeting the Core: A Simple and Effective Method to Attack RAG-based Agents via Direct LLM Manipulation
Viaarxiv icon

Precision Knowledge Editing: Enhancing Safety in Large Language Models

Add code
Oct 02, 2024
Figure 1 for Precision Knowledge Editing: Enhancing Safety in Large Language Models
Figure 2 for Precision Knowledge Editing: Enhancing Safety in Large Language Models
Viaarxiv icon