Picture for Sicheng Zhu

Sicheng Zhu

Michael Pokorny

GPT-Red: Automated Red Teaming via Self-Play at Scale

Add code
Jul 28, 2026
Viaarxiv icon

IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs

Add code
Mar 11, 2026
Viaarxiv icon

MIST-RL: Mutation-based Incremental Suite Testing via Reinforcement Learning

Add code
Mar 02, 2026
Viaarxiv icon

Weakly-supervised Contrastive Learning with Quantity Prompts for Moving Infrared Small Target Detection

Add code
Jul 03, 2025
Viaarxiv icon

Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds

Add code
Apr 02, 2025
Figure 1 for Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
Figure 2 for Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
Figure 3 for Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
Figure 4 for Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
Viaarxiv icon

PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models

Add code
Mar 10, 2025
Figure 1 for PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
Figure 2 for PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
Figure 3 for PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
Figure 4 for PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
Viaarxiv icon

Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals

Add code
Dec 15, 2024
Viaarxiv icon

AdvPrefix: An Objective for Nuanced LLM Jailbreaks

Add code
Dec 13, 2024
Figure 1 for AdvPrefix: An Objective for Nuanced LLM Jailbreaks
Figure 2 for AdvPrefix: An Objective for Nuanced LLM Jailbreaks
Figure 3 for AdvPrefix: An Objective for Nuanced LLM Jailbreaks
Figure 4 for AdvPrefix: An Objective for Nuanced LLM Jailbreaks
Viaarxiv icon

GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Add code
Oct 10, 2024
Figure 1 for GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
Figure 2 for GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
Figure 3 for GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
Figure 4 for GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
Viaarxiv icon

Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Add code
Sep 01, 2024
Figure 1 for Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Figure 2 for Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Figure 3 for Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Figure 4 for Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Viaarxiv icon