Picture for Victor Lu

Victor Lu

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

Add code
Jul 07, 2026
Viaarxiv icon

AI Benchmark Democratization and Carpentry

Add code
Dec 12, 2025
Viaarxiv icon

Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world

Add code
Oct 29, 2025
Viaarxiv icon

Risk Management for Mitigating Benchmark Failure Modes: BenchRisk

Add code
Oct 24, 2025
Viaarxiv icon

Introducing v0.5 of the AI Safety Benchmark from MLCommons

Add code
Apr 18, 2024
Figure 1 for Introducing v0.5 of the AI Safety Benchmark from MLCommons
Figure 2 for Introducing v0.5 of the AI Safety Benchmark from MLCommons
Figure 3 for Introducing v0.5 of the AI Safety Benchmark from MLCommons
Figure 4 for Introducing v0.5 of the AI Safety Benchmark from MLCommons
Viaarxiv icon