Picture for Amit LeVi

Amit LeVi

Unsupervised Features Mining via Activation Geometry

Add code
Jul 05, 2026
Viaarxiv icon

Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring

Add code
Jun 29, 2026
Viaarxiv icon

Mirage Probes: How Vision Models Fake Visual Understanding

Add code
Jun 11, 2026
Viaarxiv icon

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software

Add code
Feb 02, 2026
Viaarxiv icon

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models

Add code
Feb 01, 2026
Viaarxiv icon

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

Add code
Nov 09, 2025
Figure 1 for You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
Figure 2 for You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
Viaarxiv icon

Silenced Biases: The Dark Side LLMs Learned to Refuse

Add code
Nov 05, 2025
Viaarxiv icon