Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Douwe Kiela

I am a Strange Dataset: Metalinguistic Tests for Language Models

Jan 10, 2024

Tristan Thrush, Jared Moore, Miguel Monares, Christopher Potts, Douwe Kiela

Figure 1 for I am a Strange Dataset: Metalinguistic Tests for Language Models

Figure 2 for I am a Strange Dataset: Metalinguistic Tests for Language Models

Figure 3 for I am a Strange Dataset: Metalinguistic Tests for Language Models

Figure 4 for I am a Strange Dataset: Metalinguistic Tests for Language Models

Abstract:Statements involving metalinguistic self-reference ("This paper has six sections.") are prevalent in many domains. Can large language models (LLMs) handle such language? In this paper, we present "I am a Strange Dataset", a new dataset for addressing this question. There are two subtasks: generation and verification. In generation, models continue statements like "The penultimate word in this sentence is" (where a correct continuation is "is"). In verification, models judge the truth of statements like "The penultimate word in this sentence is sentence." (false). We also provide minimally different metalinguistic non-self-reference examples to complement the main dataset by probing for whether models can handle metalinguistic language at all. The dataset is hand-crafted by experts and validated by non-expert annotators. We test a variety of open-source LLMs (7B to 70B parameters) as well as closed-source LLMs through APIs. All models perform close to chance across both subtasks and even on the non-self-referential metalinguistic control data, though we find some steady improvement with model scale. GPT 4 is the only model to consistently do significantly better than chance, and it is still only in the 60% range, while our untrained human annotators score well in the 89-93% range. The dataset and evaluation toolkit are available at https://github.com/TristanThrush/i-am-a-strange-dataset.

Via

Access Paper or Ask Questions

Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision

Nov 25, 2023

Nicholas Lui, Bryan Chia, William Berrios, Candace Ross, Douwe Kiela

Figure 1 for Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision

Figure 2 for Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision

Figure 3 for Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision

Figure 4 for Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision

Abstract:Computer vision models have been known to encode harmful biases, leading to the potentially unfair treatment of historically marginalized groups, such as people of color. However, there remains a lack of datasets balanced along demographic traits that can be used to evaluate the downstream fairness of these models. In this work, we demonstrate that diffusion models can be leveraged to create such a dataset. We first use a diffusion model to generate a large set of images depicting various occupations. Subsequently, each image is edited using inpainting to generate multiple variants, where each variant refers to a different perceived race. Using this dataset, we benchmark several vision-language models on a multi-class occupation classification task. We find that images generated with non-Caucasian labels have a significantly higher occupation misclassification rate than images generated with Caucasian labels, and that several misclassifications are suggestive of racial biases. We measure a model's downstream fairness by computing the standard deviation in the probability of predicting the true occupation label across the different perceived identity groups. Using this fairness metric, we find significant disparities between the evaluated vision-and-language models. We hope that our work demonstrates the potential value of diffusion methods for fairness evaluations.

* The Appendix can be found at https://bit.ly/dp-appendix

Via

Access Paper or Ask Questions

FinanceBench: A New Benchmark for Financial Question Answering

Nov 20, 2023

Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, Bertie Vidgen

Figure 1 for FinanceBench: A New Benchmark for Financial Question Answering

Figure 2 for FinanceBench: A New Benchmark for Financial Question Answering

Figure 3 for FinanceBench: A New Benchmark for Financial Question Answering

Figure 4 for FinanceBench: A New Benchmark for Financial Question Answering

Abstract:FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). It comprises 10,231 questions about publicly traded companies, with corresponding answers and evidence strings. The questions in FinanceBench are ecologically valid and cover a diverse set of scenarios. They are intended to be clear-cut and straightforward to answer to serve as a minimum performance standard. We test 16 state of the art model configurations (including GPT-4-Turbo, Llama2 and Claude2, with vector stores and long context prompts) on a sample of 150 cases from FinanceBench, and manually review their answers (n=2,400). The cases are available open-source. We show that existing LLMs have clear limitations for financial QA. Notably, GPT-4-Turbo used with a retrieval system incorrectly answered or refused to answer 81% of questions. While augmentation techniques such as using longer context window to feed in relevant evidence improve performance, they are unrealistic for enterprise settings due to increased latency and cannot support larger financial documents. We find that all models examined exhibit weaknesses, such as hallucinations, that limit their suitability for use by enterprises.

* Dataset is available at: https://huggingface.co/datasets/PatronusAI/financebench

Via

Access Paper or Ask Questions

Anchor Points: Benchmarking Models with Much Fewer Examples

Sep 14, 2023

Rajan Vivek, Kawin Ethayarajh, Diyi Yang, Douwe Kiela

Figure 1 for Anchor Points: Benchmarking Models with Much Fewer Examples

Figure 2 for Anchor Points: Benchmarking Models with Much Fewer Examples

Figure 3 for Anchor Points: Benchmarking Models with Much Fewer Examples

Figure 4 for Anchor Points: Benchmarking Models with Much Fewer Examples

Abstract:Modern language models often exhibit powerful but brittle behavior, leading to the development of larger and more diverse benchmarks to reliably assess their behavior. Here, we suggest that model performance can be benchmarked and elucidated with much smaller evaluation sets. We first show that in six popular language classification benchmarks, model confidence in the correct class on many pairs of points is strongly correlated across models. We build upon this phenomenon to propose Anchor Point Selection, a technique to select small subsets of datasets that capture model behavior across the entire dataset. Anchor points reliably rank models: across 87 diverse language model-prompt pairs, evaluating models using 1-30 anchor points outperforms uniform sampling and other baselines at accurately ranking models. Moreover, just several anchor points can be used to estimate model per-class predictions on all other points in a dataset with low mean absolute error, sufficient for gauging where the model is likely to fail. Lastly, we present Anchor Point Maps for visualizing these insights and facilitating comparisons of the performance of different models on various regions within the dataset distribution.

Via

Access Paper or Ask Questions

Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

Jun 28, 2023

William Berrios, Gautam Mittal, Tristan Thrush, Douwe Kiela, Amanpreet Singh

Figure 1 for Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

Figure 2 for Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

Figure 3 for Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

Figure 4 for Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

Abstract:We propose LENS, a modular approach for tackling computer vision problems by leveraging the power of large language models (LLMs). Our system uses a language model to reason over outputs from a set of independent and highly descriptive vision modules that provide exhaustive information about an image. We evaluate the approach on pure computer vision settings such as zero- and few-shot object recognition, as well as on vision and language problems. LENS can be applied to any off-the-shelf LLM and we find that the LLMs with LENS perform highly competitively with much bigger and much more sophisticated systems, without any multimodal training whatsoever. We open-source our code at https://github.com/ContextualAI/lens and provide an interactive demo.

Via

Access Paper or Ask Questions

OBELISC: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Jun 21, 2023

Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela(+2 more)

Figure 1 for OBELISC: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Figure 2 for OBELISC: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Figure 3 for OBELISC: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Figure 4 for OBELISC: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Abstract:Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks that require reasoning over one or multiple images to generate a text. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. We introduce the OBELISC dataset, an open web-scale filtered dataset of interleaved image-text documents comprising 141 million web pages extracted from Common Crawl, 353 million associated images, and 115 billion text tokens. We describe the dataset creation process, present comprehensive filtering rules, and provide an analysis of the dataset's content. To show the viability of OBELISC, we train an 80 billion parameters vision and language model on the dataset and obtain competitive performance on various multimodal benchmarks. We release the code to reproduce the dataset along with the dataset itself.

Via

Access Paper or Ask Questions

AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages

Apr 04, 2023

Chris Chinenye Emezue, Sanchit Gandhi, Lewis Tunstall, Abubakar Abid, Josh Meyer, Quentin Lhoest, Pete Allen, Patrick Von Platen, Douwe Kiela, Yacine Jernite(+3 more)

Abstract:The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minimalist, community-driven dataset of spoken digits for African languages, currently covering 38 African languages. As a demonstration of the practical applications of AfroDigits, we conduct audio digit classification experiments on six African languages [Igbo (ibo), Yoruba (yor), Rundi (run), Oshiwambo (kua), Shona (sna), and Oromo (gax)] using the Wav2Vec2.0-Large and XLS-R models. Our experiments reveal a useful insight on the effect of mixing African speech corpora during finetuning. AfroDigits is the first published audio digit dataset for African languages and we believe it will, among other things, pave the way for Afro-centric speech applications such as the recognition of telephone numbers, and street numbers. We release the dataset and platform publicly at https://huggingface.co/datasets/chrisjay/crowd-speech-africa and https://huggingface.co/spaces/chrisjay/afro-speech respectively.

* Accepted to the AfricaNLP Workshop at ICLR 2023

Via

Access Paper or Ask Questions

Investigating Multi-source Active Learning for Natural Language Inference

Feb 14, 2023

Ard Snijders, Douwe Kiela, Katerina Margatina

Figure 1 for Investigating Multi-source Active Learning for Natural Language Inference

Figure 2 for Investigating Multi-source Active Learning for Natural Language Inference

Figure 3 for Investigating Multi-source Active Learning for Natural Language Inference

Figure 4 for Investigating Multi-source Active Learning for Natural Language Inference

Abstract:In recent years, active learning has been successfully applied to an array of NLP tasks. However, prior work often assumes that training and test data are drawn from the same distribution. This is problematic, as in real-life settings data may stem from several sources of varying relevance and quality. We show that four popular active learning schemes fail to outperform random selection when applied to unlabelled pools comprised of multiple data sources on the task of natural language inference. We reveal that uncertainty-based strategies perform poorly due to the acquisition of collective outliers, i.e., hard-to-learn instances that hamper learning and generalization. When outliers are removed, strategies are found to recover and outperform random baselines. In further analysis, we find that collective outliers vary in form between sources, and show that hard-to-learn data is not always categorically harmful. Lastly, we leverage dataset cartography to introduce difficulty-stratified testing and find that different strategies are affected differently by example learnability and difficulty.

* 23 pages. Accepted for publication at the European Chapter of the Association of Computational Linguistics (EACL) 2023

Via

Access Paper or Ask Questions

Measuring Data

Dec 09, 2022

Margaret Mitchell, Alexandra Sasha Luccioni, Nathan Lambert, Marissa Gerchick, Angelina McMillan-Major, Ezinwanne Ozoani, Nazneen Rajani, Tristan Thrush, Yacine Jernite, Douwe Kiela

Abstract:We identify the task of measuring data to quantitatively characterize the composition of machine learning data and datasets. Similar to an object's height, width, and volume, data measurements quantify different attributes of data along common dimensions that support comparison. Several lines of research have proposed what we refer to as measurements, with differing terminology; we bring some of this work together, particularly in fields of computer vision and language, and build from it to motivate measuring data as a critical component of responsible AI development. Measuring data aids in systematically building and analyzing machine learning (ML) data towards specific goals and gaining better control of what modern ML systems will learn. We conclude with a discussion of the many avenues of future work, the limitations of data measurements, and how to leverage these measurement approaches in research and practice.

Via

Access Paper or Ask Questions

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Nov 09, 2022

Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé(+380 more)

Abstract:Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed by resource-rich organizations and are frequently kept from the public. As a step towards democratizing this powerful technology, we present BLOOM, a 176B-parameter open-access language model designed and built thanks to a collaboration of hundreds of researchers. BLOOM is a decoder-only Transformer language model that was trained on the ROOTS corpus, a dataset comprising hundreds of sources in 46 natural and 13 programming languages (59 in total). We find that BLOOM achieves competitive performance on a wide variety of benchmarks, with stronger results after undergoing multitask prompted finetuning. To facilitate future research and applications using LLMs, we publicly release our models and code under the Responsible AI License.

Via

Access Paper or Ask Questions