Picture for Minlie Huang

Minlie Huang

EJ

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

Add code
Jul 01, 2026
Viaarxiv icon

SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

Add code
Jun 02, 2026
Viaarxiv icon

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

Add code
Jun 02, 2026
Viaarxiv icon

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

Add code
May 28, 2026
Viaarxiv icon

You Live More Than Once: Towards Hierarchical Skill Meta-Evolving

Add code
May 27, 2026
Viaarxiv icon

EVA: Editing for Versatile Alignment against Jailbreaks

Add code
May 14, 2026
Viaarxiv icon

HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing

Add code
Apr 21, 2026
Viaarxiv icon

LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety

Add code
Apr 13, 2026
Viaarxiv icon

SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions

Add code
Mar 24, 2026
Viaarxiv icon

IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation

Add code
Mar 05, 2026
Viaarxiv icon