Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

James Lucassen

BashArena: A Control Setting for Highly Privileged AI Agents

Dec 17, 2025

Adam Kaufman, James Lucassen, Tyler Tracy, Cody Rushing, Aryan Bhatt

Figure 1 for BashArena: A Control Setting for Highly Privileged AI Agents

Figure 2 for BashArena: A Control Setting for Highly Privileged AI Agents

Figure 3 for BashArena: A Control Setting for Highly Privileged AI Agents

Figure 4 for BashArena: A Control Setting for Highly Privileged AI Agents

Abstract:Future AI agents might run autonomously with elevated privileges. If these agents are misaligned, they might abuse these privileges to cause serious damage. The field of AI control develops techniques that make it harder for misaligned AIs to cause such damage, while preserving their usefulness. We introduce BashArena, a setting for studying AI control techniques in security-critical environments. BashArena contains 637 Linux system administration and infrastructure engineering tasks in complex, realistic environments, along with four sabotage objectives (execute malware, exfiltrate secrets, escalate privileges, and disable firewall) for a red team to target. We evaluate multiple frontier LLMs on their ability to complete tasks, perform sabotage undetected, and detect sabotage attempts. Claude Sonnet 4.5 successfully executes sabotage while evading monitoring by GPT-4.1 mini 26% of the time, at 4% trajectory-wise FPR. Our findings provide a baseline for designing more effective control protocols in BashArena. We release the dataset as a ControlArena setting and share our task generation pipeline.

* The task generation pipeline can be found here: https://github.com/redwoodresearch/basharena_public

Via

Access Paper or Ask Questions

Evaluating Stability of Unreflective Alignment

Aug 27, 2024

James Lucassen, Mark Henry, Philippa Wright, Owen Yeung

Abstract:Many theoretical obstacles to AI alignment are consequences of reflective stability - the problem of designing alignment mechanisms that the AI would not disable if given the option. However, problems stemming from reflective stability are not obviously present in current LLMs, leading to disagreement over whether they will need to be solved to enable safe delegation of cognitive labor. In this paper, we propose Counterfactual Priority Change (CPC) destabilization as a mechanism by which reflective stability problems may arise in future LLMs. We describe two risk factors for CPC-destabilization: 1) CPC-based stepping back and 2) preference instability. We develop preliminary evaluations for each of these risk factors, and apply them to frontier LLMs. Our findings indicate that in current LLMs, increased scale and capability are associated with increases in both CPC-based stepping back and preference instability, suggesting that CPC-destabilization may cause reflective stability problems in future LLMs.

Via

Access Paper or Ask Questions