Picture for Cindy Wu

Cindy Wu

Targeted Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Add code
Jul 22, 2024
Figure 1 for Targeted Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Figure 2 for Targeted Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Figure 3 for Targeted Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Figure 4 for Targeted Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Viaarxiv icon

Using Degeneracy in the Loss Landscape for Mechanistic Interpretability

Add code
May 17, 2024
Figure 1 for Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
Viaarxiv icon