Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models

Jun 16, 2022

Maribeth Rauh, John Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel(+2 more)

Figure 1 for Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models

Figure 2 for Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models

Figure 3 for Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models

Share this with someone who'll enjoy it:

Abstract:Large language models produce human-like text that drive a growing number of applications. However, recent literature and, increasingly, real world observations, have demonstrated that these models can generate language that is toxic, biased, untruthful or otherwise harmful. Though work to evaluate language model harms is under way, translating foresight about which harms may arise into rigorous benchmarks is not straightforward. To facilitate this translation, we outline six ways of characterizing harmful text which merit explicit consideration when designing new benchmarks. We then use these characteristics as a lens to identify trends and gaps in existing benchmarks. Finally, we apply them in a case study of the Perspective API, a toxicity classifier that is widely used in harm benchmarks. Our characteristics provide one piece of the bridge that translates between foresight and effective evaluation.

* 9 pages plus appendix

View paper on

OpenReview

Share this with someone who'll enjoy it:

Title:Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models

Paper and Code