Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Sebastian Farquhar

An Approach to Technical AGI Safety and Security

Apr 02, 2025

Rohin Shah, Alex Irpan, Alexander Matt Turner, Anna Wang, Arthur Conmy, David Lindner, Jonah Brown-Cohen, Lewis Ho, Neel Nanda, Raluca Ada Popa(+20 more)

Abstract:Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough to significantly harm humanity. We identify four areas of risk: misuse, misalignment, mistakes, and structural risks. Of these, we focus on technical approaches to misuse and misalignment. For misuse, our strategy aims to prevent threat actors from accessing dangerous capabilities, by proactively identifying dangerous capabilities, and implementing robust security, access restrictions, monitoring, and model safety mitigations. To address misalignment, we outline two lines of defense. First, model-level mitigations such as amplified oversight and robust training can help to build an aligned model. Second, system-level security measures such as monitoring and access control can mitigate harm even if the model is misaligned. Techniques from interpretability, uncertainty estimation, and safer design patterns can enhance the effectiveness of these mitigations. Finally, we briefly outline how these ingredients could be combined to produce safety cases for AGI systems.

Via

Access Paper or Ask Questions

Do Multilingual LLMs Think In English?

Feb 21, 2025

Lisa Schut, Yarin Gal, Sebastian Farquhar

Abstract:Large language models (LLMs) have multilingual capabilities and can solve tasks across various languages. However, we show that current LLMs make key decisions in a representation space closest to English, regardless of their input and output languages. Exploring the internal representations with a logit lens for sentences in French, German, Dutch, and Mandarin, we show that the LLM first emits representations close to English for semantically-loaded words before translating them into the target language. We further show that activation steering in these LLMs is more effective when the steering vectors are computed in English rather than in the language of the inputs and outputs. This suggests that multilingual LLMs perform key reasoning steps in a representation that is heavily shaped by English in a way that is not transparent to system users.

* Main paper 9 pages; including appendix 48 pages

Via

Access Paper or Ask Questions

Holistic Safety and Responsibility Evaluations of Advanced AI Models

Apr 22, 2024

Laura Weidinger, Joslyn Barnhart, Jenny Brennan, Christina Butterfield, Susie Young, Will Hawkins, Lisa Anne Hendricks, Ramona Comanescu, Oscar Chang, Mikel Rodriguez(+9 more)

Abstract:Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models, we innovated on and applied a broad set of approaches to safety evaluation. In this report, we summarise and share elements of our evolving approach as well as lessons learned for a broad audience. Key lessons learned include: First, theoretical underpinnings and frameworks are invaluable to organise the breadth of risk domains, modalities, forms, metrics, and goals. Second, theory and practice of safety evaluation development each benefit from collaboration to clarify goals, methods and challenges, and facilitate the transfer of insights between different stakeholders and disciplines. Third, similar key methods, lessons, and institutions apply across the range of concerns in responsibility and safety - including established and emerging harms. For this reason it is important that a wide range of actors working on safety evaluation and safety research communities work together to develop, refine and implement novel evaluation approaches and best practices, rather than operating in silos. The report concludes with outlining the clear need to rapidly advance the science of evaluations, to integrate new evaluations into the development and governance of AI, to establish scientifically-grounded norms and standards, and to promote a robust evaluation ecosystem.

* 10 pages excluding bibliography

Via

Access Paper or Ask Questions

Evaluating Frontier Models for Dangerous Capabilities

Mar 20, 2024

Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson(+17 more)

Figure 1 for Evaluating Frontier Models for Dangerous Capabilities

Figure 2 for Evaluating Frontier Models for Dangerous Capabilities

Figure 3 for Evaluating Frontier Models for Dangerous Capabilities

Figure 4 for Evaluating Frontier Models for Dangerous Capabilities

Abstract:To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models. Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs. Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.

Via

Access Paper or Ask Questions

Challenges with unsupervised LLM knowledge discovery

Dec 18, 2023

Sebastian Farquhar, Vikrant Varma, Zachary Kenton, Johannes Gasteiger, Vladimir Mikulik, Rohin Shah

Figure 1 for Challenges with unsupervised LLM knowledge discovery

Figure 2 for Challenges with unsupervised LLM knowledge discovery

Figure 3 for Challenges with unsupervised LLM knowledge discovery

Figure 4 for Challenges with unsupervised LLM knowledge discovery

Abstract:We show that existing unsupervised methods on large language model (LLM) activations do not discover knowledge -- instead they seem to discover whatever feature of the activations is most prominent. The idea behind unsupervised knowledge elicitation is that knowledge satisfies a consistency structure, which can be used to discover knowledge. We first prove theoretically that arbitrary features (not just knowledge) satisfy the consistency structure of a particular leading unsupervised knowledge-elicitation method, contrast-consistent search (Burns et al. - arXiv:2212.03827). We then present a series of experiments showing settings in which unsupervised methods result in classifiers that do not predict knowledge, but instead predict a different prominent feature. We conclude that existing unsupervised methods for discovering latent knowledge are insufficient, and we contribute sanity checks to apply to evaluating future knowledge elicitation methods. Conceptually, we hypothesise that the identification issues explored here, e.g. distinguishing a model's knowledge from that of a simulated character's, will persist for future unsupervised methods.

* 12 pages (38 including references and appendices). First three authors equal contribution, randomised order

Via

Access Paper or Ask Questions

Model evaluation for extreme risks

May 24, 2023

Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt(+11 more)

Figure 1 for Model evaluation for extreme risks

Figure 2 for Model evaluation for extreme risks

Figure 3 for Model evaluation for extreme risks

Figure 4 for Model evaluation for extreme risks

Abstract:Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further progress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is critical for addressing extreme risks. Developers must be able to identify dangerous capabilities (through "dangerous capability evaluations") and the propensity of models to apply their capabilities for harm (through "alignment evaluations"). These evaluations will become critical for keeping policymakers and other stakeholders informed, and for making responsible decisions about model training, deployment, and security.

Via

Access Paper or Ask Questions

Prediction-Oriented Bayesian Active Learning

Apr 17, 2023

Freddie Bickford Smith, Andreas Kirsch, Sebastian Farquhar, Yarin Gal, Adam Foster, Tom Rainforth

Figure 1 for Prediction-Oriented Bayesian Active Learning

Figure 2 for Prediction-Oriented Bayesian Active Learning

Figure 3 for Prediction-Oriented Bayesian Active Learning

Figure 4 for Prediction-Oriented Bayesian Active Learning

Abstract:Information-theoretic approaches to active learning have traditionally focused on maximising the information gathered about the model parameters, most commonly by optimising the BALD score. We highlight that this can be suboptimal from the perspective of predictive performance. For example, BALD lacks a notion of an input distribution and so is prone to prioritise data of limited relevance. To address this we propose the expected predictive information gain (EPIG), an acquisition function that measures information gain in the space of predictions rather than parameters. We find that using EPIG leads to stronger predictive performance compared with BALD across a range of datasets and models, and thus provides an appealing drop-in replacement.

* Published at AISTATS 2023

Via

Access Paper or Ask Questions

Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Feb 21, 2023

Lorenz Kuhn, Yarin Gal, Sebastian Farquhar

Figure 1 for Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Figure 2 for Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Figure 3 for Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Figure 4 for Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Abstract:We introduce a method to measure uncertainty in large language models. For tasks like question answering, it is essential to know when we can trust the natural language outputs of foundation models. We show that measuring uncertainty in natural language is challenging because of "semantic equivalence" -- different sentences can mean the same thing. To overcome these challenges we introduce semantic entropy -- an entropy which incorporates linguistic invariances created by shared meanings. Our method is unsupervised, uses only a single model, and requires no modifications to off-the-shelf language models. In comprehensive ablation studies we show that the semantic entropy is more predictive of model accuracy on question answering data sets than comparable baselines.

Via

Access Paper or Ask Questions

CLAM: Selective Clarification for Ambiguous Questions with Large Language Models

Dec 15, 2022

Lorenz Kuhn, Yarin Gal, Sebastian Farquhar

Figure 1 for CLAM: Selective Clarification for Ambiguous Questions with Large Language Models

Figure 2 for CLAM: Selective Clarification for Ambiguous Questions with Large Language Models

Figure 3 for CLAM: Selective Clarification for Ambiguous Questions with Large Language Models

Figure 4 for CLAM: Selective Clarification for Ambiguous Questions with Large Language Models

Abstract:State-of-the-art language models are often accurate on many question-answering benchmarks with well-defined questions. Yet, in real settings questions are often unanswerable without asking the user for clarifying information. We show that current SotA models often do not ask the user for clarification when presented with imprecise questions and instead provide incorrect answers or "hallucinate". To address this, we introduce CLAM, a framework that first uses the model to detect ambiguous questions, and if an ambiguous question is detected, prompts the model to ask the user for clarification. Furthermore, we show how to construct a scalable and cost-effective automatic evaluation protocol using an oracle language model with privileged information to provide clarifying information. We show that our method achieves a 20.15 percentage point accuracy improvement over SotA on a novel ambiguous question-answering answering data set derived from TriviaQA.

Via

Access Paper or Ask Questions

Understanding Approximation for Bayesian Inference in Neural Networks

Nov 11, 2022

Sebastian Farquhar

Figure 1 for Understanding Approximation for Bayesian Inference in Neural Networks

Figure 2 for Understanding Approximation for Bayesian Inference in Neural Networks

Figure 3 for Understanding Approximation for Bayesian Inference in Neural Networks

Figure 4 for Understanding Approximation for Bayesian Inference in Neural Networks

Abstract:Bayesian inference has theoretical attractions as a principled framework for reasoning about beliefs. However, the motivations of Bayesian inference which claim it to be the only 'rational' kind of reasoning do not apply in practice. They create a binary split in which all approximate inference is equally 'irrational'. Instead, we should ask ourselves how to define a spectrum of more- and less-rational reasoning that explains why we might prefer one Bayesian approximation to another. I explore approximate inference in Bayesian neural networks and consider the unintended interactions between the probabilistic model, approximating distribution, optimization algorithm, and dataset. The complexity of these interactions highlights the difficulty of any strategy for evaluating Bayesian approximations which focuses entirely on the method, outside the context of specific datasets and decision-problems. For given applications, the expected utility of the approximate posterior can measure inference quality. To assess a model's ability to incorporate different parts of the Bayesian framework we can identify desirable characteristic behaviours of Bayesian reasoning and pick decision-problems that make heavy use of those behaviours. Here, we use continual learning (testing the ability to update sequentially) and active learning (testing the ability to represent credence). But existing continual and active learning set-ups pose challenges that have nothing to do with posterior quality which can distort their ability to evaluate Bayesian approximations. These unrelated challenges can be removed or reduced, allowing better evaluation of approximate inference methods.

* Accepted as a thesis satisfying the requirements of a D.Phil at the Universty of Oxford

Via

Access Paper or Ask Questions