Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Lars Kai Hansen

Missing-Data-Induced Phase Transitions in Spectral PLS for Multimodal Learning

Jan 29, 2026

Anders Gjølbye, Ida Kargaard, Emma Kargaard, Lars Kai Hansen

Abstract:Partial Least Squares (PLS) learns shared structure from paired data via the top singular vectors of the empirical cross-covariance (PLS-SVD), but multimodal datasets often have missing entries in both views. We study PLS-SVD under independent entry-wise missing-completely-at-random masking in a proportional high-dimensional spiked model. After appropriate normalization, the masked cross-covariance behaves like a spiked rectangular random matrix whose effective signal strength is attenuated by $\sqrtρ$, where $ρ$ is the joint entry retention probability. As a result, PLS-SVD exhibits a sharp BBP-type phase transition: below a critical signal-to-noise threshold the leading singular vectors are asymptotically uninformative, while above it they achieve nontrivial alignment with the latent shared directions, with closed-form asymptotic overlap formulas. Simulations and semi-synthetic multimodal experiments corroborate the predicted phase diagram and recovery curves across aspect ratios, signal strengths, and missingness levels.

* Preprint

Via

Access Paper or Ask Questions

Large Vision Models Can Solve Mental Rotation Problems

Sep 18, 2025

Sebastian Ray Mason, Anders Gjølbye, Phillip Chavarria Højbjerg, Lenka Tětková, Lars Kai Hansen

Abstract:Mental rotation is a key test of spatial reasoning in humans and has been central to understanding how perception supports cognition. Despite the success of modern vision transformers, it is still unclear how well these models develop similar abilities. In this work, we present a systematic evaluation of ViT, CLIP, DINOv2, and DINOv3 across a range of mental-rotation tasks, from simple block structures similar to those used by Shepard and Metzler to study human cognition, to more complex block figures, three types of text, and photo-realistic objects. By probing model representations layer by layer, we examine where and how these networks succeed. We find that i) self-supervised ViTs capture geometric structure better than supervised ViTs; ii) intermediate layers perform better than final layers; iii) task difficulty increases with rotation complexity and occlusion, mirroring human reaction times and suggesting similar constraints in embedding space representations.

Via

Access Paper or Ask Questions

Minimizing False-Positive Attributions in Explanations of Non-Linear Models

May 16, 2025

Anders Gjølbye, Stefan Haufe, Lars Kai Hansen

Abstract:Suppressor variables can influence model predictions without being dependent on the target outcome and they pose a significant challenge for Explainable AI (XAI) methods. These variables may cause false-positive feature attributions, undermining the utility of explanations. Although effective remedies exist for linear models, their extension to non-linear models and to instance-based explanations has remained limited. We introduce PatternLocal, a novel XAI technique that addresses this gap. PatternLocal begins with a locally linear surrogate, e.g. LIME, KernelSHAP, or gradient-based methods, and transforms the resulting discriminative model weights into a generative representation, thereby suppressing the influence of suppressor variables while preserving local fidelity. In extensive hyperparameter optimization on the XAI-TRIS benchmark, PatternLocal consistently outperformed other XAI methods and reduced false-positive attributions when explaining non-linear tasks, thereby enabling more reliable and actionable insights.

* Preprint. Under review

Via

Access Paper or Ask Questions

From Colors to Classes: Emergence of Concepts in Vision Transformers

Mar 31, 2025

Teresa Dorszewski, Lenka Tětková, Robert Jenssen, Lars Kai Hansen, Kristoffer Knutsen Wickstrøm

Figure 1 for From Colors to Classes: Emergence of Concepts in Vision Transformers

Figure 2 for From Colors to Classes: Emergence of Concepts in Vision Transformers

Figure 3 for From Colors to Classes: Emergence of Concepts in Vision Transformers

Figure 4 for From Colors to Classes: Emergence of Concepts in Vision Transformers

Abstract:Vision Transformers (ViTs) are increasingly utilized in various computer vision tasks due to their powerful representation capabilities. However, it remains understudied how ViTs process information layer by layer. Numerous studies have shown that convolutional neural networks (CNNs) extract features of increasing complexity throughout their layers, which is crucial for tasks like domain adaptation and transfer learning. ViTs, lacking the same inductive biases as CNNs, can potentially learn global dependencies from the first layers due to their attention mechanisms. Given the increasing importance of ViTs in computer vision, there is a need to improve the layer-wise understanding of ViTs. In this work, we present a novel, layer-wise analysis of concepts encoded in state-of-the-art ViTs using neuron labeling. Our findings reveal that ViTs encode concepts with increasing complexity throughout the network. Early layers primarily encode basic features such as colors and textures, while later layers represent more specific classes, including objects and animals. As the complexity of encoded concepts increases, the number of concepts represented in each layer also rises, reflecting a more diverse and specific set of features. Additionally, different pretraining strategies influence the quantity and category of encoded concepts, with finetuning to specific downstream tasks generally reducing the number of encoded concepts and shifting the concepts to more relevant categories.

* Preprint. Accepted at The 3rd World Conference on eXplainable Artificial Intelligence

Via

Access Paper or Ask Questions

Danoliteracy of Generative, Large Language Models

Oct 30, 2024

Søren Vejlgaard Holm, Lars Kai Hansen, Martin Carsten Nielsen

Figure 1 for Danoliteracy of Generative, Large Language Models

Figure 2 for Danoliteracy of Generative, Large Language Models

Figure 3 for Danoliteracy of Generative, Large Language Models

Figure 4 for Danoliteracy of Generative, Large Language Models

Abstract:The language technology moonshot moment of Generative, Large Language Models (GLLMs) was not limited to English: These models brought a surge of technological applications, investments and hype to low-resource languages as well. However, the capabilities of these models in languages such as Danish were until recently difficult to verify beyond qualitative demonstrations due to a lack of applicable evaluation corpora. We present a GLLM benchmark to evaluate Danoliteracy, a measure of Danish language and cultural competency, across eight diverse scenarios such Danish citizenship tests and abstractive social media question answering. This limited-size benchmark is found to produce a robust ranking that correlates to human feedback at $\rho \sim 0.8$ with GPT-4 and Claude Opus models achieving the highest rankings. Analyzing these model results across scenarios, we find one strong underlying factor explaining $95\%$ of scenario performance variance for GLLMs in Danish, suggesting a $g$ factor of model consistency in language adaption.

* 16 pages, 13 figures, submitted to: NoDaLiDa/Baltic-HLT 2025

Via

Access Paper or Ask Questions

BiSSL: Bilevel Optimization for Self-Supervised Pre-Training and Fine-Tuning

Oct 03, 2024

Gustav Wagner Zakarias, Lars Kai Hansen, Zheng-Hua Tan

Figure 1 for BiSSL: Bilevel Optimization for Self-Supervised Pre-Training and Fine-Tuning

Figure 2 for BiSSL: Bilevel Optimization for Self-Supervised Pre-Training and Fine-Tuning

Figure 3 for BiSSL: Bilevel Optimization for Self-Supervised Pre-Training and Fine-Tuning

Figure 4 for BiSSL: Bilevel Optimization for Self-Supervised Pre-Training and Fine-Tuning

Abstract:In this work, we present BiSSL, a first-of-its-kind training framework that introduces bilevel optimization to enhance the alignment between the pretext pre-training and downstream fine-tuning stages in self-supervised learning. BiSSL formulates the pretext and downstream task objectives as the lower- and upper-level objectives in a bilevel optimization problem and serves as an intermediate training stage within the self-supervised learning pipeline. By more explicitly modeling the interdependence of these training stages, BiSSL facilitates enhanced information sharing between them, ultimately leading to a backbone parameter initialization that is better suited for the downstream task. We propose a training algorithm that alternates between optimizing the two objectives defined in BiSSL. Using a ResNet-18 backbone pre-trained with SimCLR on the STL10 dataset, we demonstrate that our proposed framework consistently achieves improved or competitive classification accuracies across various downstream image classification datasets compared to the conventional self-supervised learning pipeline. Qualitative analyses of the backbone features further suggest that BiSSL enhances the alignment of downstream features in the backbone prior to fine-tuning.

Via

Access Paper or Ask Questions

Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks

Sep 10, 2024

Teresa Dorszewski, Lenka Tětková, Lorenz Linhardt, Lars Kai Hansen

Figure 1 for Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks

Figure 2 for Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks

Figure 3 for Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks

Figure 4 for Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks

Abstract:Understanding how neural networks align with human cognitive processes is a crucial step toward developing more interpretable and reliable AI systems. Motivated by theories of human cognition, this study examines the relationship between \emph{convexity} in neural network representations and \emph{human-machine alignment} based on behavioral data. We identify a correlation between these two dimensions in pretrained and fine-tuned vision transformer models. Our findings suggest that the convex regions formed in latent spaces of neural networks to some extent align with human-defined categories and reflect the similarity relations humans use in cognitive tasks. While optimizing for alignment generally enhances convexity, increasing convexity through fine-tuning yields inconsistent effects on alignment, which suggests a complex relationship between the two. This study presents a first step toward understanding the relationship between the convexity of latent representations and human-machine alignment.

* First two authors contributed equally

Via

Access Paper or Ask Questions

Convexity-based Pruning of Speech Representation Models

Aug 16, 2024

Teresa Dorszewski, Lenka Tětková, Lars Kai Hansen

Figure 1 for Convexity-based Pruning of Speech Representation Models

Figure 2 for Convexity-based Pruning of Speech Representation Models

Figure 3 for Convexity-based Pruning of Speech Representation Models

Abstract:Speech representation models based on the transformer architecture and trained by self-supervised learning have shown great promise for solving tasks such as speech and speaker recognition, keyword spotting, emotion detection, and more. Typically, it is found that larger models lead to better performance. However, the significant computational effort involved in such large transformer systems is a challenge for embedded and real-world applications. Recent work has shown that there is significant redundancy in the transformer models for NLP and massive layer pruning is feasible (Sajjad et al., 2023). Here, we investigate layer pruning in audio models. We base the pruning decision on a convexity criterion. Convexity of classification regions has recently been proposed as an indicator of subsequent fine-tuning performance in a range of application domains, including NLP and audio. In empirical investigations, we find a massive reduction in the computational effort with no loss of performance or even improvements in certain cases.

Via

Access Paper or Ask Questions

SPEED: Scalable Preprocessing of EEG Data for Self-Supervised Learning

Aug 15, 2024

Anders Gjølbye, Lina Skerath, William Lehn-Schiøler, Nicolas Langer, Lars Kai Hansen

Abstract:Electroencephalography (EEG) research typically focuses on tasks with narrowly defined objectives, but recent studies are expanding into the use of unlabeled data within larger models, aiming for a broader range of applications. This addresses a critical challenge in EEG research. For example, Kostas et al. (2021) show that self-supervised learning (SSL) outperforms traditional supervised methods. Given the high noise levels in EEG data, we argue that further improvements are possible with additional preprocessing. Current preprocessing methods often fail to efficiently manage the large data volumes required for SSL, due to their lack of optimization, reliance on subjective manual corrections, and validation processes or inflexible protocols that limit SSL. We propose a Python-based EEG preprocessing pipeline optimized for self-supervised learning, designed to efficiently process large-scale data. This optimization not only stabilizes self-supervised training but also enhances performance on downstream tasks compared to training with raw data.

* To appear in proceedings of 2024 IEEE International workshop on Machine Learning for Signal Processing

Via

Access Paper or Ask Questions

Challenges in explaining deep learning models for data with biological variation

Jun 14, 2024

Lenka Tětková, Erik Schou Dreier, Robin Malm, Lars Kai Hansen

Figure 1 for Challenges in explaining deep learning models for data with biological variation

Figure 2 for Challenges in explaining deep learning models for data with biological variation

Figure 3 for Challenges in explaining deep learning models for data with biological variation

Figure 4 for Challenges in explaining deep learning models for data with biological variation

Abstract:Much machine learning research progress is based on developing models and evaluating them on a benchmark dataset (e.g., ImageNet for images). However, applying such benchmark-successful methods to real-world data often does not work as expected. This is particularly the case for biological data where we expect variability at multiple time and spatial scales. In this work, we are using grain data and the goal is to detect diseases and damages. Pink fusarium, skinned grains, and other diseases and damages are key factors in setting the price of grains or excluding dangerous grains from food production. Apart from challenges stemming from differences of the data from the standard toy datasets, we also present challenges that need to be overcome when explaining deep learning models. For example, explainability methods have many hyperparameters that can give different results, and the ones published in the papers do not work on dissimilar images. Other challenges are more general: problems with visualization of the explanations and their comparison since the magnitudes of their values differ from method to method. An open fundamental question also is: How to evaluate explanations? It is a non-trivial task because the "ground truth" is usually missing or ill-defined. Also, human annotators may create what they think is an explanation of the task at hand, yet the machine learning model might solve it in a different and perhaps counter-intuitive way. We discuss several of these challenges and evaluate various post-hoc explainability methods on grain data. We focus on robustness, quality of explanations, and similarity to particular "ground truth" annotations made by experts. The goal is to find the methods that overall perform well and could be used in this challenging task. We hope the proposed pipeline will be used as a framework for evaluating explainability methods in specific use cases.

Via

Access Paper or Ask Questions