Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Pierre Ablin

Ecole normale supérieure, Paris, France

MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations

Jan 13, 2025

Ambroise Heurtebise, Omar Chehab, Pierre Ablin, Alexandre Gramfort

Figure 1 for MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations

Figure 2 for MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations

Figure 3 for MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations

Figure 4 for MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations

Abstract:Machine learning techniques in multi-view settings face significant challenges, particularly when integrating heterogeneous data, aligning feature spaces, and managing view-specific biases. These issues are prominent in neuroscience, where data from multiple subjects exposed to the same stimuli are analyzed to uncover brain activity dynamics. In magnetoencephalography (MEG), where signals are captured at the scalp level, estimating the brain's underlying sources is crucial, especially in group studies where sources are assumed to be similar for all subjects. Common methods, such as Multi-View Independent Component Analysis (MVICA), assume identical sources across subjects, but this assumption is often too restrictive due to individual variability and age-related changes. Multi-View Independent Component Analysis with Delays (MVICAD) addresses this by allowing sources to differ up to a temporal delay. However, temporal dilation effects, particularly in auditory stimuli, are common in brain dynamics, making the estimation of time delays alone insufficient. To address this, we propose Multi-View Independent Component Analysis with Delays and Dilations (MVICAD2), which allows sources to differ across subjects in both temporal delays and dilations. We present a model with identifiable sources, derive an approximation of its likelihood in closed form, and use regularization and optimization techniques to enhance performance. Through simulations, we demonstrate that MVICAD2 outperforms existing multi-view ICA methods. We further validate its effectiveness using the Cam-CAN dataset, and showing how delays and dilations are related to aging.

* 19 pages, 8 figures

Via

Access Paper or Ask Questions

Sparse Repellency for Shielded Generation in Text-to-image Diffusion Models

Oct 10, 2024

Michael Kirchhof, James Thornton, Pierre Ablin, Louis Béthune, Eugene Ndiaye, Marco Cuturi

Figure 1 for Sparse Repellency for Shielded Generation in Text-to-image Diffusion Models

Figure 2 for Sparse Repellency for Shielded Generation in Text-to-image Diffusion Models

Figure 3 for Sparse Repellency for Shielded Generation in Text-to-image Diffusion Models

Figure 4 for Sparse Repellency for Shielded Generation in Text-to-image Diffusion Models

Abstract:The increased adoption of diffusion models in text-to-image generation has triggered concerns on their reliability. Such models are now closely scrutinized under the lens of various metrics, notably calibration, fairness, or compute efficiency. We focus in this work on two issues that arise when deploying these models: a lack of diversity when prompting images, and a tendency to recreate images from the training set. To solve both problems, we propose a method that coaxes the sampled trajectories of pretrained diffusion models to land on images that fall outside of a reference set. We achieve this by adding repellency terms to the diffusion SDE throughout the generation trajectory, which are triggered whenever the path is expected to land too closely to an image in the shielded reference set. Our method is sparse in the sense that these repellency terms are zero and inactive most of the time, and even more so towards the end of the generation trajectory. Our method, named SPELL for sparse repellency, can be used either with a static reference set that contains protected images, or dynamically, by updating the set at each timestep with the expected images concurrently generated within a batch. We show that adding SPELL to popular diffusion models improves their diversity while impacting their FID only marginally, and performs comparatively better than other recent training-free diversity methods. We also demonstrate how SPELL can ensure a shielded generation away from a very large set of protected images by considering all 1.2M images from ImageNet as the protected set.

Via

Access Paper or Ask Questions

Dynamic Gradient Alignment for Online Data Mixing

Oct 03, 2024

Simin Fan, David Grangier, Pierre Ablin

Figure 1 for Dynamic Gradient Alignment for Online Data Mixing

Figure 2 for Dynamic Gradient Alignment for Online Data Mixing

Figure 3 for Dynamic Gradient Alignment for Online Data Mixing

Figure 4 for Dynamic Gradient Alignment for Online Data Mixing

Abstract:The composition of training data mixtures is critical for effectively training large language models (LLMs), as it directly impacts their performance on downstream tasks. Our goal is to identify an optimal data mixture to specialize an LLM for a specific task with access to only a few examples. Traditional approaches to this problem include ad-hoc reweighting methods, importance sampling, and gradient alignment techniques. This paper focuses on gradient alignment and introduces Dynamic Gradient Alignment (DGA), a scalable online gradient alignment algorithm. DGA dynamically estimates the pre-training data mixture on which the models' gradients align as well as possible with those of the model on the specific task. DGA is the first gradient alignment approach that incurs minimal overhead compared to standard pre-training and outputs a competitive model, eliminating the need for retraining the model. Experimentally, we demonstrate significant improvements over importance sampling in two key scenarios: (i) when the pre-training set is small and importance sampling overfits due to limited data; and (ii) when there is insufficient specialized data, trapping importance sampling on narrow pockets of data. Our findings underscore the effectiveness of gradient alignment methods in optimizing training data mixtures, particularly in data-constrained environments, and offer a practical solution for enhancing LLM performance on specific tasks with limited data availability.

Via

Access Paper or Ask Questions

Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Sep 06, 2024

Jason Ramapuram, Federico Danieli, Eeshan Dhekane, Floris Weers, Dan Busbridge, Pierre Ablin, Tatiana Likhomanenko, Jagrit Digani, Zijin Gu, Amitis Shidani(+1 more)

Figure 1 for Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Figure 2 for Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Figure 3 for Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Figure 4 for Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Abstract:Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are typically obtained as the softmax of dot products between keys and queries. Recent work has explored alternatives to softmax attention in transformers, such as ReLU and sigmoid activations. In this work, we revisit sigmoid attention and conduct an in-depth theoretical and empirical analysis. Theoretically, we prove that transformers with sigmoid attention are universal function approximators and benefit from improved regularity compared to softmax attention. Through detailed empirical analysis, we identify stabilization of large initial attention norms during the early stages of training as a crucial factor for the successful training of models with sigmoid attention, outperforming prior attempts. We also introduce FLASHSIGMOID, a hardware-aware and memory-efficient implementation of sigmoid attention yielding a 17% inference kernel speed-up over FLASHATTENTION2 on H100 GPUs. Experiments across language, vision, and speech show that properly normalized sigmoid attention matches the strong performance of softmax attention on a wide range of domains and scales, which previous attempts at sigmoid attention were unable to fully achieve. Our work unifies prior art and establishes best practices for sigmoid attention as a drop-in softmax replacement in transformers.

Via

Access Paper or Ask Questions

The AdEMAMix Optimizer: Better, Faster, Older

Sep 05, 2024

Matteo Pagliardini, Pierre Ablin, David Grangier

Figure 1 for The AdEMAMix Optimizer: Better, Faster, Older

Figure 2 for The AdEMAMix Optimizer: Better, Faster, Older

Figure 3 for The AdEMAMix Optimizer: Better, Faster, Older

Figure 4 for The AdEMAMix Optimizer: Better, Faster, Older

Abstract:Momentum based optimizers are central to a wide range of machine learning applications. These typically rely on an Exponential Moving Average (EMA) of gradients, which decays exponentially the present contribution of older gradients. This accounts for gradients being local linear approximations which lose their relevance as the iterate moves along the loss landscape. This work questions the use of a single EMA to accumulate past gradients and empirically demonstrates how this choice can be sub-optimal: a single EMA cannot simultaneously give a high weight to the immediate past, and a non-negligible weight to older gradients. Building on this observation, we propose AdEMAMix, a simple modification of the Adam optimizer with a mixture of two EMAs to better take advantage of past gradients. Our experiments on language modeling and image classification show -- quite surprisingly -- that gradients can stay relevant for tens of thousands of steps. They help to converge faster, and often to lower minima: e.g., a $1.3$B parameter AdEMAMix LLM trained on $101$B tokens performs comparably to an AdamW model trained on $197$B tokens ($+95\%$). Moreover, our method significantly slows-down model forgetting during training. Our work motivates further exploration of different types of functions to leverage past gradients, beyond EMAs.

* 38 pages, 27 figures

Via

Access Paper or Ask Questions

Optimization without retraction on the random generalized Stiefel manifold

May 02, 2024

Simon Vary, Pierre Ablin, Bin Gao, P. -A. Absil

Figure 1 for Optimization without retraction on the random generalized Stiefel manifold

Figure 2 for Optimization without retraction on the random generalized Stiefel manifold

Figure 3 for Optimization without retraction on the random generalized Stiefel manifold

Figure 4 for Optimization without retraction on the random generalized Stiefel manifold

Abstract:Optimization over the set of matrices that satisfy $X^\top B X = I_p$, referred to as the generalized Stiefel manifold, appears in many applications involving sampled covariance matrices such as canonical correlation analysis (CCA), independent component analysis (ICA), and the generalized eigenvalue problem (GEVP). Solving these problems is typically done by iterative methods, such as Riemannian approaches, which require a computationally expensive eigenvalue decomposition involving fully formed $B$. We propose a cheap stochastic iterative method that solves the optimization problem while having access only to a random estimate of the feasible set. Our method does not enforce the constraint in every iteration exactly, but instead it produces iterations that converge to a critical point on the generalized Stiefel manifold defined in expectation. The method has lower per-iteration cost, requires only matrix multiplications, and has the same convergence rates as its Riemannian counterparts involving the full matrix $B$. Experiments demonstrate its effectiveness in various machine learning applications involving generalized orthogonality constraints, including CCA, ICA, and GEVP.

* 21 pages, 10 figures

Via

Access Paper or Ask Questions

Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization

Feb 26, 2024

Zhenzhang Ye, Gabriel Peyré, Daniel Cremers, Pierre Ablin

Abstract:Bilevel optimization aims to optimize an outer objective function that depends on the solution to an inner optimization problem. It is routinely used in Machine Learning, notably for hyperparameter tuning. The conventional method to compute the so-called hypergradient of the outer problem is to use the Implicit Function Theorem (IFT). As a function of the error of the inner problem resolution, we study the error of the IFT method. We analyze two strategies to reduce this error: preconditioning the IFT formula and reparameterizing the inner problem. We give a detailed account of the impact of these two modifications on the error, highlighting the role played by higher-order derivatives of the functionals at stake. Our theoretical findings explain when super efficiency, namely reaching an error on the hypergradient that depends quadratically on the error on the inner problem, is achievable and compare the two approaches when this is impossible. Numerical evaluations on hyperparameter tuning for regression problems substantiate our theoretical findings.

* Accepted in AISTATS 2024

Via

Access Paper or Ask Questions

Careful with that Scalpel: Improving Gradient Surgery with an EMA

Feb 05, 2024

Yu-Guan Hsieh, James Thornton, Eugene Ndiaye, Michal Klein, Marco Cuturi, Pierre Ablin

Abstract:Beyond minimizing a single training loss, many deep learning estimation pipelines rely on an auxiliary objective to quantify and encourage desirable properties of the model (e.g. performance on another dataset, robustness, agreement with a prior). Although the simplest approach to incorporating an auxiliary loss is to sum it with the training loss as a regularizer, recent works have shown that one can improve performance by blending the gradients beyond a simple sum; this is known as gradient surgery. We cast the problem as a constrained minimization problem where the auxiliary objective is minimized among the set of minimizers of the training loss. To solve this bilevel problem, we follow a parameter update direction that combines the training loss gradient and the orthogonal projection of the auxiliary gradient to the training gradient. In a setting where gradients come from mini-batches, we explain how, using a moving average of the training loss gradients, we can carefully maintain this critical orthogonality property. We demonstrate that our method, Bloop, can lead to much better performances on NLP and vision experiments than other gradient surgery methods without EMA.

Via

Access Paper or Ask Questions

Specialized Language Models with Cheap Inference from Limited Domain Data

Feb 02, 2024

David Grangier, Angelos Katharopoulos, Pierre Ablin, Awni Hannun

Figure 1 for Specialized Language Models with Cheap Inference from Limited Domain Data

Figure 2 for Specialized Language Models with Cheap Inference from Limited Domain Data

Figure 3 for Specialized Language Models with Cheap Inference from Limited Domain Data

Figure 4 for Specialized Language Models with Cheap Inference from Limited Domain Data

Abstract:Large language models have emerged as a versatile tool but are challenging to apply to tasks lacking large inference budgets and large in-domain training sets. This work formalizes these constraints and distinguishes four important variables: the pretraining budget (for training before the target domain is known), the specialization budget (for training after the target domain is known), the inference budget, and the in-domain training set size. Across these settings, we compare different approaches from the machine learning literature. Limited by inference cost, we find better alternatives to the standard practice of training very large vanilla transformer models. In particular, we show that hyper-networks and mixture of experts have better perplexity for large pretraining budgets, while small models trained on importance sampled datasets are attractive for large specialization budgets.

Via

Access Paper or Ask Questions

Understanding the Regularity of Self-Attention with Optimal Transport

Dec 22, 2023

Valérie Castin, Pierre Ablin, Gabriel Peyré

Figure 1 for Understanding the Regularity of Self-Attention with Optimal Transport

Figure 2 for Understanding the Regularity of Self-Attention with Optimal Transport

Figure 3 for Understanding the Regularity of Self-Attention with Optimal Transport

Figure 4 for Understanding the Regularity of Self-Attention with Optimal Transport

Abstract:Transformers and their multi-head attention mechanism have completely changed the machine learning landscape in just a few years, by outperforming state-of-art models in a wide range of domains. Still, little is known about their robustness from a theoretical perspective. We tackle this problem by studying the local Lipschitz constant of self-attention, that provides an attack-agnostic way of measuring the robustness of a neural network. We adopt a measure-theoretic framework, by viewing inputs as probability measures equipped with the Wasserstein distance. This allows us to generalize attention to inputs of infinite length, and to derive an upper bound and a lower bound on the Lipschitz constant of self-attention on compact sets. The lower bound significantly improves prior results, and grows more than exponentially with the radius of the compact set, which rules out the possibility of obtaining robustness guarantees without any additional constraint on the input space. Our results also point out that measures with a high local Lipschitz constant are typically made of a few diracs, with a very unbalanced distribution of mass. Finally, we analyze the stability of self-attention under perturbations that change the number of tokens, which appears to be a natural question in the measure-theoretic framework. In particular, we show that for some inputs, attacks that duplicate tokens before perturbing them are more efficient than attacks that simply move tokens. We call this phenomenon mass splitting.

Via

Access Paper or Ask Questions