Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Nathan Godey

The State-Prediction Separation Hypothesis

Jul 01, 2026

Giovanni Monea, Nathan Godey, Kianté Brantley, Yoav Artzi

Abstract:Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the \emph{state-prediction separation hypothesis}: disentangling the two roles yields better language modeling performance. We design a Transformer variant that uses two computation streams to separate the two functions, and conduct pretraining experiments across various scales. Our experiments show that state-prediction separation consistently offers better data and compute efficiencies, improving validation loss and outperforming standard Transformers by 2--3 percentage points on average on downstream tasks. We also conduct extensive empirical analysis that rules out potential confounders and demonstrates the fundamental difference in the gradients our design entails.

* Preprint

Via

Access Paper or Ask Questions

Lost in Backpropagation: The LM Head is a Gradient Bottleneck

Mar 10, 2026

Nathan Godey, Yoav Artzi

Abstract:The last layer of neural language models (LMs) projects output features of dimension $D$ to logits in dimension $V$, the size of the vocabulary, where usually $D \ll V$. This mismatch is known to raise risks of limited expressivity in neural LMs, creating a so-called softmax bottleneck. We show the softmax bottleneck is not only an expressivity bottleneck but also an optimization bottleneck. Backpropagating $V$-dimensional gradients through a rank-$D$ linear layer induces unavoidable compression, which alters the training feedback provided to the vast majority of the parameters. We present a theoretical analysis of this phenomenon and measure empirically that 95-99% of the gradient norm is suppressed by the output layer, resulting in vastly suboptimal update directions. We conduct controlled pretraining experiments showing that the gradient bottleneck makes trivial patterns unlearnable, and drastically affects the training dynamics of LLMs. We argue that this inherent flaw contributes to training inefficiencies at scale independently of the model architecture, and raises the need for new LM head designs.

Via

Access Paper or Ask Questions

Gaperon: A Peppered English-French Generative Language Model Suite

Oct 29, 2025

Nathan Godey, Wissam Antoun, Rian Touchent, Rachel Bawden, Éric de la Clergerie, Benoît Sagot, Djamé Seddah

Abstract:We release Gaperon, a fully open suite of French-English-coding language models designed to advance transparency and reproducibility in large-scale model training. The Gaperon family includes 1.5B, 8B, and 24B parameter models trained on 2-4 trillion tokens, released with all elements of the training pipeline: French and English datasets filtered with a neural quality classifier, an efficient data curation and training framework, and hundreds of intermediate checkpoints. Through this work, we study how data filtering and contamination interact to shape both benchmark and generative performance. We find that filtering for linguistic quality enhances text fluency and coherence but yields subpar benchmark results, and that late deliberate contamination -- continuing training on data mixes that include test sets -- recovers competitive scores while only reasonably harming generation quality. We discuss how usual neural filtering can unintentionally amplify benchmark leakage. To support further research, we also introduce harmless data poisoning during pretraining, providing a realistic testbed for safety studies. By openly releasing all models, datasets, code, and checkpoints, Gaperon establishes a reproducible foundation for exploring the trade-offs between data curation, evaluation, safety, and openness in multilingual language model development.

Via

Access Paper or Ask Questions

Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression

Mar 04, 2025

Nathan Godey, Alessio Devoto, Yu Zhao, Simone Scardapane, Pasquale Minervini, Éric de la Clergerie, Benoît Sagot

Abstract:Autoregressive language models rely on a Key-Value (KV) Cache, which avoids re-computing past hidden states during generation, making it faster. As model sizes and context lengths grow, the KV Cache becomes a significant memory bottleneck, which calls for compression methods that limit its size during generation. In this paper, we discover surprising properties of Query (Q) and Key (K) vectors that allow us to efficiently approximate attention scores without computing the attention maps. We propose Q-Filters, a training-free KV Cache compression method that filters out less crucial Key-Value pairs based on a single context-agnostic projection. Contrarily to many alternatives, Q-Filters is compatible with FlashAttention, as it does not require direct access to attention weights. Experimental results in long-context settings demonstrate that Q-Filters is competitive with attention-based compression methods such as SnapKV in retrieval tasks while consistently outperforming efficient compression schemes such as Streaming-LLM in generation setups. Notably, Q-Filters achieves a 99% accuracy in the needle-in-a-haystack task with a x32 compression level while reducing the generation perplexity drop by up to 65% in text generation compared to Streaming-LLM.

Via

Access Paper or Ask Questions

Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck

Apr 11, 2024

Nathan Godey, Éric de la Clergerie, Benoît Sagot

Abstract:Recent advances in language modeling consist in pretraining highly parameterized neural networks on extremely large web-mined text corpora. Training and inference with such models can be costly in practice, which incentivizes the use of smaller counterparts. However, it has been observed that smaller models can suffer from saturation, characterized as a drop in performance at some advanced point in training followed by a plateau. In this paper, we find that such saturation can be explained by a mismatch between the hidden dimension of smaller models and the high rank of the target contextual probability distribution. This mismatch affects the performance of the linear prediction head used in such models through the well-known softmax bottleneck phenomenon. We measure the effect of the softmax bottleneck in various settings and find that models based on less than 1000 hidden dimensions tend to adopt degenerate latent representations in late pretraining, which leads to reduced evaluation performance.

Via

Access Paper or Ask Questions

On the Scaling Laws of Geographical Representation in Language Models

Mar 04, 2024

Nathan Godey, Éric de la Clergerie, Benoît Sagot

Figure 1 for On the Scaling Laws of Geographical Representation in Language Models

Figure 2 for On the Scaling Laws of Geographical Representation in Language Models

Figure 3 for On the Scaling Laws of Geographical Representation in Language Models

Figure 4 for On the Scaling Laws of Geographical Representation in Language Models

Abstract:Language models have long been shown to embed geographical information in their hidden representations. This line of work has recently been revisited by extending this result to Large Language Models (LLMs). In this paper, we propose to fill the gap between well-established and recent literature by observing how geographical knowledge evolves when scaling language models. We show that geographical knowledge is observable even for tiny models, and that it scales consistently as we increase the model size. Notably, we observe that larger language models cannot mitigate the geographical bias that is inherent to the training data.

* Accepted at LREC-COLING 2024

Via

Access Paper or Ask Questions

Anisotropy Is Inherent to Self-Attention in Transformers

Jan 24, 2024

Nathan Godey, Éric de la Clergerie, Benoît Sagot

Figure 1 for Anisotropy Is Inherent to Self-Attention in Transformers

Figure 2 for Anisotropy Is Inherent to Self-Attention in Transformers

Figure 3 for Anisotropy Is Inherent to Self-Attention in Transformers

Figure 4 for Anisotropy Is Inherent to Self-Attention in Transformers

Abstract:The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers. In NLP, it takes the form of anisotropy, a singular property of hidden representations which makes them unexpectedly close to each other in terms of angular distance (cosine-similarity). Some recent works tend to show that anisotropy is a consequence of optimizing the cross-entropy loss on long-tailed distributions of tokens. We show in this paper that anisotropy can also be observed empirically in language models with specific objectives that should not suffer directly from the same consequences. We also show that the anisotropy problem extends to Transformers trained on other modalities. Our observations suggest that anisotropy is actually inherent to Transformers-based models.

* Proceedings of EACL 2024. A previous version of the paper, published as arXiv:2306.07656, was presented at ACL-SRW 2023 (non-archival)

Via

Access Paper or Ask Questions

Headless Language Models: Learning without Predicting with Contrastive Weight Tying

Sep 15, 2023

Nathan Godey, Éric de la Clergerie, Benoît Sagot

Figure 1 for Headless Language Models: Learning without Predicting with Contrastive Weight Tying

Figure 2 for Headless Language Models: Learning without Predicting with Contrastive Weight Tying

Figure 3 for Headless Language Models: Learning without Predicting with Contrastive Weight Tying

Figure 4 for Headless Language Models: Learning without Predicting with Contrastive Weight Tying

Abstract:Self-supervised pre-training of language models usually consists in predicting probability distributions over extensive token vocabularies. In this study, we propose an innovative method that shifts away from probability prediction and instead focuses on reconstructing input embeddings in a contrastive fashion via Constrastive Weight Tying (CWT). We apply this approach to pretrain Headless Language Models in both monolingual and multilingual contexts. Our method offers practical advantages, substantially reducing training computational requirements by up to 20 times, while simultaneously enhancing downstream performance and data efficiency. We observe a significant +1.6 GLUE score increase and a notable +2.7 LAMBADA accuracy improvement compared to classical LMs within similar compute budgets.

Via

Access Paper or Ask Questions

Is Anisotropy Inherent to Transformers?

Jun 13, 2023

Nathan Godey, Éric de la Clergerie, Benoît Sagot

Abstract:The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers. In NLP, it takes the form of anisotropy, a singular property of hidden representations which makes them unexpectedly close to each other in terms of angular distance (cosine-similarity). Some recent works tend to show that anisotropy is a consequence of optimizing the cross-entropy loss on long-tailed distributions of tokens. We show in this paper that anisotropy can also be observed empirically in language models with specific objectives that should not suffer directly from the same consequences. We also show that the anisotropy problem extends to Transformers trained on other modalities. Our observations tend to demonstrate that anisotropy might actually be inherent to Transformers-based models.

* ACL-SRW 2023 (Poster)

Via

Access Paper or Ask Questions

MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling

Dec 14, 2022

Nathan Godey, Roman Castagné, Éric de la Clergerie, Benoît Sagot

Figure 1 for MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling

Figure 2 for MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling

Figure 3 for MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling

Figure 4 for MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling

Abstract:Static subword tokenization algorithms have been an essential component of recent works on language modeling. However, their static nature results in important flaws that degrade the models' downstream performance and robustness. In this work, we propose MANTa, a Module for Adaptive Neural TokenizAtion. MANTa is a differentiable tokenizer trained end-to-end with the language model. The resulting system offers a trade-off between the expressiveness of byte-level models and the speed of models trained using subword tokenization. In addition, our tokenizer is highly explainable since it produces an explicit segmentation of sequences into blocks. We evaluate our pre-trained model on several English datasets from different domains as well as on synthetic noise. We find that MANTa improves robustness to character perturbations and out-of-domain data. We then show that MANTa performs comparably to other models on the general-domain GLUE benchmark. Finally, we show that it is considerably faster than strictly byte-level models.

* EMNLP 2022 Findings (https://aclanthology.org/2022.findings-emnlp.207/)

Via

Access Paper or Ask Questions