Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Henryk Michalewski

A Simple, Yet Effective Approach to Finding Biases in Code Generation

Oct 31, 2022

Spyridon Mouselinos, Mateusz Malinowski, Henryk Michalewski

Figure 1 for A Simple, Yet Effective Approach to Finding Biases in Code Generation

Figure 2 for A Simple, Yet Effective Approach to Finding Biases in Code Generation

Figure 3 for A Simple, Yet Effective Approach to Finding Biases in Code Generation

Figure 4 for A Simple, Yet Effective Approach to Finding Biases in Code Generation

Abstract:Recently, scores of high-performing code generation systems have surfaced. As has become a popular choice in many domains, code generation is often approached using large language models as a core, trained under the masked or causal language modeling schema. This work shows that current code generation systems exhibit biases inherited from large language model backbones, which might leak into generated code under specific circumstances. To investigate the effect, we propose a framework that automatically removes hints and exposes various biases that these code generation models use. We apply our framework to three coding challenges and test it across top-performing coding generation models. Our experiments reveal biases towards specific prompt structure and exploitation of keywords during code generation. Finally, we demonstrate how to use our framework as a data transformation technique, which we find a promising direction toward more robust code generation.

* Preprint

Via

Access Paper or Ask Questions

Language Model Cascades

Jul 28, 2022

David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A. Saurous, Jascha Sohl-dickstein(+2 more)

Abstract:Prompted models have demonstrated impressive few-shot learning abilities. Repeated interactions at test-time with a single model, or the composition of multiple models together, further expands capabilities. These compositions are probabilistic models, and may be expressed in the language of graphical models with random variables whose values are complex data types such as strings. Cases with control flow and dynamic structure require techniques from probabilistic programming, which allow implementing disparate model structures and inference strategies in a unified language. We formalize several existing techniques from this perspective, including scratchpads / chain of thought, verifiers, STaR, selection-inference, and tool use. We refer to the resulting programs as language model cascades.

* Presented as spotlight at the Beyond Bases workshop at ICML 2022 (https://beyond-bayes.github.io)

Via

Access Paper or Ask Questions

Solving Quantitative Reasoning Problems with Language Models

Jul 01, 2022

Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo(+4 more)

Figure 1 for Solving Quantitative Reasoning Problems with Language Models

Figure 2 for Solving Quantitative Reasoning Problems with Language Models

Figure 3 for Solving Quantitative Reasoning Problems with Language Models

Figure 4 for Solving Quantitative Reasoning Problems with Language Models

Abstract:Language models have achieved remarkable performance on a wide range of tasks that require natural language understanding. Nevertheless, state-of-the-art models have generally struggled with tasks that require quantitative reasoning, such as solving mathematics, science, and engineering problems at the college level. To help close this gap, we introduce Minerva, a large language model pretrained on general natural language data and further trained on technical content. The model achieves state-of-the-art performance on technical benchmarks without the use of external tools. We also evaluate our model on over two hundred undergraduate-level problems in physics, biology, chemistry, economics, and other sciences that require quantitative reasoning, and find that the model can correctly answer nearly a third of them.

* 12 pages, 5 figures + references and appendices

Via

Access Paper or Ask Questions

Multi-Game Decision Transformers

May 30, 2022

Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee, Daniel Freeman, Winnie Xu, Sergio Guadarrama, Ian Fischer, Eric Jang, Henryk Michalewski(+1 more)

Figure 1 for Multi-Game Decision Transformers

Figure 2 for Multi-Game Decision Transformers

Figure 3 for Multi-Game Decision Transformers

Figure 4 for Multi-Game Decision Transformers

Abstract:A longstanding goal of the field of AI is a strategy for compiling diverse experience into a highly capable, generalist agent. In the subfields of vision and language, this was largely achieved by scaling up transformer-based models and training them on large, diverse datasets. Motivated by this progress, we investigate whether the same strategy can be used to produce generalist reinforcement learning agents. Specifically, we show that a single transformer-based model - with a single set of weights - trained purely offline can play a suite of up to 46 Atari games simultaneously at close-to-human performance. When trained and evaluated appropriately, we find that the same trends observed in language and vision hold, including scaling of performance with model size and rapid adaptation to new games via fine-tuning. We compare several approaches in this multi-game setting, such as online and offline RL methods and behavioral cloning, and find that our Multi-Game Decision Transformer models offer the best scalability and performance. We release the pre-trained models and code to encourage further research in this direction. Additional information, videos and code can be seen at: sites.google.com/view/multi-game-transformers

Via

Access Paper or Ask Questions

PaLM: Scaling Language Modeling with Pathways

Apr 19, 2022

Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann(+57 more)

Figure 1 for PaLM: Scaling Language Modeling with Pathways

Figure 2 for PaLM: Scaling Language Modeling with Pathways

Figure 3 for PaLM: Scaling Language Modeling with Pathways

Figure 4 for PaLM: Scaling Language Modeling with Pathways

Abstract:Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM. We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.

Via

Access Paper or Ask Questions

Measuring CLEVRness: Blackbox testing of Visual Reasoning Models

Feb 28, 2022

Spyridon Mouselinos, Henryk Michalewski, Mateusz Malinowski

Figure 1 for Measuring CLEVRness: Blackbox testing of Visual Reasoning Models

Figure 2 for Measuring CLEVRness: Blackbox testing of Visual Reasoning Models

Figure 3 for Measuring CLEVRness: Blackbox testing of Visual Reasoning Models

Figure 4 for Measuring CLEVRness: Blackbox testing of Visual Reasoning Models

Abstract:How can we measure the reasoning capabilities of intelligence systems? Visual question answering provides a convenient framework for testing the model's abilities by interrogating the model through questions about the scene. However, despite scores of various visual QA datasets and architectures, which sometimes yield even a super-human performance, the question of whether those architectures can actually reason remains open to debate. To answer this, we extend the visual question answering framework and propose the following behavioral test in the form of a two-player game. We consider black-box neural models of CLEVR. These models are trained on a diagnostic dataset benchmarking reasoning. Next, we train an adversarial player that re-configures the scene to fool the CLEVR model. We show that CLEVR models, which otherwise could perform at a human level, can easily be fooled by our agent. Our results put in doubt whether data-driven approaches can do reasoning without exploiting the numerous biases that are often present in those datasets. Finally, we also propose a controlled experiment measuring the efficiency of such models to learn and perform reasoning.

* ICLR 2022

Via

Access Paper or Ask Questions

Show Your Work: Scratchpads for Intermediate Computation with Language Models

Nov 30, 2021

Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan(+2 more)

Figure 1 for Show Your Work: Scratchpads for Intermediate Computation with Language Models

Figure 2 for Show Your Work: Scratchpads for Intermediate Computation with Language Models

Figure 3 for Show Your Work: Scratchpads for Intermediate Computation with Language Models

Figure 4 for Show Your Work: Scratchpads for Intermediate Computation with Language Models

Abstract:Large pre-trained language models perform remarkably well on tasks that can be done "in one pass", such as generating realistic text or synthesizing computer programs. However, they struggle with tasks that require unbounded multi-step computation, such as adding integers or executing programs. Surprisingly, we find that these same models are able to perform complex multi-step computations -- even in the few-shot regime -- when asked to perform the operation "step by step", showing the results of intermediate computations. In particular, we train transformers to perform multi-step computations by asking them to emit intermediate computation steps into a "scratchpad". On a series of increasingly complex tasks ranging from long addition to the execution of arbitrary programs, we show that scratchpads dramatically improve the ability of language models to perform multi-step computations.

Via

Access Paper or Ask Questions

Sparse is Enough in Scaling Transformers

Nov 24, 2021

Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Łukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, Jonni Kanerva

Figure 1 for Sparse is Enough in Scaling Transformers

Figure 2 for Sparse is Enough in Scaling Transformers

Figure 3 for Sparse is Enough in Scaling Transformers

Figure 4 for Sparse is Enough in Scaling Transformers

Abstract:Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study sparse variants for all layers in the Transformer and propose Scaling Transformers, a family of next generation Transformer models that use sparse layers to scale efficiently and perform unbatched decoding much faster than the standard Transformer as we scale up the model size. Surprisingly, the sparse layers are enough to obtain the same perplexity as the standard Transformer with the same number of parameters. We also integrate with prior sparsity approaches to attention and enable fast inference on long sequences even with limited memory. This results in performance competitive to the state-of-the-art on long text summarization.

* NeurIPS 2021

Via

Access Paper or Ask Questions

Off-Policy Correction For Multi-Agent Reinforcement Learning

Nov 22, 2021

Michał Zawalski, Błażej Osiński, Henryk Michalewski, Piotr Miłoś

Figure 1 for Off-Policy Correction For Multi-Agent Reinforcement Learning

Figure 2 for Off-Policy Correction For Multi-Agent Reinforcement Learning

Figure 3 for Off-Policy Correction For Multi-Agent Reinforcement Learning

Figure 4 for Off-Policy Correction For Multi-Agent Reinforcement Learning

Abstract:Multi-agent reinforcement learning (MARL) provides a framework for problems involving multiple interacting agents. Despite apparent similarity to the single-agent case, multi-agent problems are often harder to train and analyze theoretically. In this work, we propose MA-Trace, a new on-policy actor-critic algorithm, which extends V-Trace to the MARL setting. The key advantage of our algorithm is its high scalability in a multi-worker setting. To this end, MA-Trace utilizes importance sampling as an off-policy correction method, which allows distributing the computations with no impact on the quality of training. Furthermore, our algorithm is theoretically grounded - we prove a fixed-point theorem that guarantees convergence. We evaluate the algorithm extensively on the StarCraft Multi-Agent Challenge, a standard benchmark for multi-agent algorithms. MA-Trace achieves high performance on all its tasks and exceeds state-of-the-art results on some of them.

Via

Access Paper or Ask Questions

Hierarchical Transformers Are More Efficient Language Models

Oct 26, 2021

Piotr Nawrot, Szymon Tworkowski, Michał Tyrolski, Łukasz Kaiser, Yuhuai Wu, Christian Szegedy, Henryk Michalewski

Figure 1 for Hierarchical Transformers Are More Efficient Language Models

Figure 2 for Hierarchical Transformers Are More Efficient Language Models

Figure 3 for Hierarchical Transformers Are More Efficient Language Models

Figure 4 for Hierarchical Transformers Are More Efficient Language Models

Abstract:Transformer models yield impressive results on many NLP and sequence modeling tasks. Remarkably, Transformers can handle long sequences which allows them to produce long coherent outputs: full paragraphs produced by GPT-3 or well-structured images produced by DALL-E. These large language models are impressive but also very inefficient and costly, which limits their applications and accessibility. We postulate that having an explicit hierarchical architecture is the key to Transformers that efficiently handle long sequences. To verify this claim, we first study different ways to downsample and upsample activations in Transformers so as to make them hierarchical. We use the best performing upsampling and downsampling layers to create Hourglass - a hierarchical Transformer language model. Hourglass improves upon the Transformer baseline given the same amount of computation and can yield the same results as Transformers more efficiently. In particular, Hourglass sets new state-of-the-art for Transformer models on the ImageNet32 generation task and improves language modeling efficiency on the widely studied enwik8 benchmark.

Via

Access Paper or Ask Questions