Picture for Ali Emami

Ali Emami

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

Add code
Aug 18, 2026
Viaarxiv icon

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

Add code
Jul 16, 2026
Viaarxiv icon

LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

Add code
Jul 07, 2026
Viaarxiv icon

Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance

Add code
Jul 06, 2026
Viaarxiv icon

Agents' Last Exam

Add code
Jun 03, 2026
Viaarxiv icon

Memory Dial: A Training Framework for Controllable Memorization in Language Models

Add code
Apr 06, 2026
Viaarxiv icon

Reasoning Traces Shape Outputs but Models Won't Say So

Add code
Mar 21, 2026
Viaarxiv icon

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

Add code
Feb 13, 2026
Viaarxiv icon

Common to Whom? Regional Cultural Commonsense and LLM Bias in India

Add code
Jan 22, 2026
Viaarxiv icon

The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts

Add code
Oct 23, 2025
Figure 1 for The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
Figure 2 for The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
Figure 3 for The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
Figure 4 for The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
Viaarxiv icon