Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Punya Syon Pandey

SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests

Oct 06, 2025

Punya Syon Pandey, Hai Son Le, Devansh Bhardwaj, Rada Mihalcea, Zhijing Jin

Figure 1 for SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests

Figure 2 for SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests

Figure 3 for SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests

Figure 4 for SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests

Abstract:Large language models (LLMs) are increasingly deployed in contexts where their failures can have direct sociopolitical consequences. Yet, existing safety benchmarks rarely test vulnerabilities in domains such as political manipulation, propaganda and disinformation generation, or surveillance and information control. We introduce SocialHarmBench, a dataset of 585 prompts spanning 7 sociopolitical categories and 34 countries, designed to surface where LLMs most acutely fail in politically charged contexts. Our evaluations reveal several shortcomings: open-weight models exhibit high vulnerability to harmful compliance, with Mistral-7B reaching attack success rates as high as 97% to 98% in domains such as historical revisionism, propaganda, and political manipulation. Moreover, temporal and geographic analyses show that LLMs are most fragile when confronted with 21st-century or pre-20th-century contexts, and when responding to prompts tied to regions such as Latin America, the USA, and the UK. These findings demonstrate that current safeguards fail to generalize to high-stakes sociopolitical settings, exposing systematic biases and raising concerns about the reliability of LLMs in preserving human rights and democratic values. We share the SocialHarmBench benchmark at https://huggingface.co/datasets/psyonp/SocialHarmBench.

Via

Access Paper or Ask Questions

Accidental Misalignment: Fine-Tuning Language Models Induces Unexpected Vulnerability

May 22, 2025

Punya Syon Pandey, Samuel Simko, Kellin Pelrine, Zhijing Jin

Figure 1 for Accidental Misalignment: Fine-Tuning Language Models Induces Unexpected Vulnerability

Figure 2 for Accidental Misalignment: Fine-Tuning Language Models Induces Unexpected Vulnerability

Figure 3 for Accidental Misalignment: Fine-Tuning Language Models Induces Unexpected Vulnerability

Figure 4 for Accidental Misalignment: Fine-Tuning Language Models Induces Unexpected Vulnerability

Abstract:As large language models gain popularity, their vulnerability to adversarial attacks remains a primary concern. While fine-tuning models on domain-specific datasets is often employed to improve model performance, it can introduce vulnerabilities within the underlying model. In this work, we investigate Accidental Misalignment, unexpected vulnerabilities arising from characteristics of fine-tuning data. We begin by identifying potential correlation factors such as linguistic features, semantic similarity, and toxicity within our experimental datasets. We then evaluate the adversarial performance of these fine-tuned models and assess how dataset factors correlate with attack success rates. Lastly, we explore potential causal links, offering new insights into adversarial defense strategies and highlighting the crucial role of dataset design in preserving model alignment. Our code is available at https://github.com/psyonp/accidental_misalignment.

Via

Access Paper or Ask Questions