Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Luccioni

What's in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus

May 06, 2021
Alexandra, Luccioni, Joseph D. Viviano

Figure 1 for What's in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus

Figure 2 for What's in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus

Figure 3 for What's in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus

Figure 4 for What's in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus

Whereas much of the success of the current generation of neural language models has been driven by increasingly large training corpora, relatively little research has been dedicated to analyzing these massive sources of textual data. In this exploratory analysis, we delve deeper into the Common Crawl, a colossal web corpus that is extensively used for training language models. We find that it contains a significant amount of undesirable content, including hate speech and sexually explicit content, even after filtering procedures. We conclude with a discussion of the potential impacts of this content on language models and call for more mindful approach to corpus collection and analysis.

* 4 pages, 1 figure, 3 tables. Published as a main conference paper at ACL-IJCNLP 2021, submission #87. Code available at https://github.com/josephdviviano/whatsinthebox

Via

Access Paper or Ask Questions

Mapping the Landscape of Artificial Intelligence Applications against COVID-19

Mar 25, 2020
Joseph Bullock, Alexandra, Luccioni, Katherine Hoffmann Pham, Cynthia Sin Nga Lam, Miguel Luengo-Oroz

Figure 1 for Mapping the Landscape of Artificial Intelligence Applications against COVID-19

COVID-19, the disease caused by the SARS-CoV-2 virus, has been declared a pandemic by the World Health Organization, with over 294,000 cases as of March 22nd 2020. In this review, we present an overview of recent studies using Machine Learning and, more broadly, Artificial Intelligence, to tackle many aspects of the COVID-19 crisis at different scales including molecular, medical and epidemiological applications. We finish with a discussion of promising future directions of research and the tools and resources needed to facilitate AI research.

* 14 pages

Via

Access Paper or Ask Questions