Picture for Twm Stone

Twm Stone

LangChoiceBench: Measuring and Explaining Programming-Language Choice in LLMs

Add code
Aug 06, 2026
Viaarxiv icon

Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

Add code
Jun 05, 2026
Viaarxiv icon

CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Add code
Mar 21, 2025
Figure 1 for CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Figure 2 for CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Figure 3 for CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Figure 4 for CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Viaarxiv icon