Picture for Peter Kirgis

Peter Kirgis

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Add code
Jul 29, 2026
Viaarxiv icon

Life After Benchmark Saturation: A Case Study of CORE-Bench

Add code
Jun 23, 2026
Viaarxiv icon

Open-World Evaluations for Measuring Frontier AI Capabilities

Add code
May 19, 2026
Viaarxiv icon

Towards a Science of AI Agent Reliability

Add code
Feb 18, 2026
Viaarxiv icon

Differences in the Moral Foundations of Large Language Models

Add code
Nov 14, 2025
Figure 1 for Differences in the Moral Foundations of Large Language Models
Figure 2 for Differences in the Moral Foundations of Large Language Models
Figure 3 for Differences in the Moral Foundations of Large Language Models
Figure 4 for Differences in the Moral Foundations of Large Language Models
Viaarxiv icon