Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Dorde Popovic

aiXamine: Simplified LLM Safety and Security

Apr 23, 2025

Fatih Deniz, Dorde Popovic, Yazan Boshmaf, Euisuh Jeong, Minhaj Ahmad, Sanjay Chawla, Issa Khalil

Figure 1 for aiXamine: Simplified LLM Safety and Security

Figure 2 for aiXamine: Simplified LLM Safety and Security

Figure 3 for aiXamine: Simplified LLM Safety and Security

Figure 4 for aiXamine: Simplified LLM Safety and Security

Abstract:Evaluating Large Language Models (LLMs) for safety and security remains a complex task, often requiring users to navigate a fragmented landscape of ad hoc benchmarks, datasets, metrics, and reporting formats. To address this challenge, we present aiXamine, a comprehensive black-box evaluation platform for LLM safety and security. aiXamine integrates over 40 tests (i.e., benchmarks) organized into eight key services targeting specific dimensions of safety and security: adversarial robustness, code security, fairness and bias, hallucination, model and data privacy, out-of-distribution (OOD) robustness, over-refusal, and safety alignment. The platform aggregates the evaluation results into a single detailed report per model, providing a detailed breakdown of model performance, test examples, and rich visualizations. We used aiXamine to assess over 50 publicly available and proprietary LLMs, conducting over 2K examinations. Our findings reveal notable vulnerabilities in leading models, including susceptibility to adversarial attacks in OpenAI's GPT-4o, biased outputs in xAI's Grok-3, and privacy weaknesses in Google's Gemini 2.0. Additionally, we observe that open-source models can match or exceed proprietary models in specific services such as safety alignment, fairness and bias, and OOD robustness. Finally, we identify trade-offs between distillation strategies, model size, training methods, and architectural choices.

Via

Access Paper or Ask Questions

aiXamine: LLM Safety and Security Simplified

Apr 21, 2025

Fatih Deniz, Dorde Popovic, Yazan Boshmaf, Euisuh Jeong, Minhaj Ahmad, Sanjay Chawla, Issa Khalil

Figure 1 for aiXamine: LLM Safety and Security Simplified

Figure 2 for aiXamine: LLM Safety and Security Simplified

Figure 3 for aiXamine: LLM Safety and Security Simplified

Figure 4 for aiXamine: LLM Safety and Security Simplified

Via

Access Paper or Ask Questions

DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data

Mar 27, 2025

Dorde Popovic, Amin Sadeghi, Ting Yu, Sanjay Chawla, Issa Khalil

Figure 1 for DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data

Figure 2 for DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data

Figure 3 for DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data

Figure 4 for DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data

Abstract:Backdoor attacks are among the most effective, practical, and stealthy attacks in deep learning. In this paper, we consider a practical scenario where a developer obtains a deep model from a third party and uses it as part of a safety-critical system. The developer wants to inspect the model for potential backdoors prior to system deployment. We find that most existing detection techniques make assumptions that are not applicable to this scenario. In this paper, we present a novel framework for detecting backdoors under realistic restrictions. We generate candidate triggers by deductively searching over the space of possible triggers. We construct and optimize a smoothed version of Attack Success Rate as our search objective. Starting from a broad class of template attacks and just using the forward pass of a deep model, we reverse engineer the backdoor attack. We conduct extensive evaluation on a wide range of attacks, models, and datasets, with our technique performing almost perfectly across these settings.

Via

Access Paper or Ask Questions