Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Michalis Vazirgiannis

Ecole Polytechnique, AUEB

LLM as a Broken Telephone: Iterative Generation Distorts Information

Feb 27, 2025

Amr Mohamed, Mingmeng Geng, Michalis Vazirgiannis, Guokan Shang

Figure 1 for LLM as a Broken Telephone: Iterative Generation Distorts Information

Figure 2 for LLM as a Broken Telephone: Iterative Generation Distorts Information

Figure 3 for LLM as a Broken Telephone: Iterative Generation Distorts Information

Figure 4 for LLM as a Broken Telephone: Iterative Generation Distorts Information

Abstract:As large language models are increasingly responsible for online content, concerns arise about the impact of repeatedly processing their own outputs. Inspired by the "broken telephone" effect in chained human communication, this study investigates whether LLMs similarly distort information through iterative generation. Through translation-based experiments, we find that distortion accumulates over time, influenced by language choice and chain complexity. While degradation is inevitable, it can be mitigated through strategic prompting techniques. These findings contribute to discussions on the long-term effects of AI-mediated information propagation, raising important questions about the reliability of LLM-generated content in iterative workflows.

Via

Access Paper or Ask Questions

Obtaining Example-Based Explanations from Deep Neural Networks

Feb 27, 2025

Genghua Dong, Henrik Boström, Michalis Vazirgiannis, Roman Bresson

Figure 1 for Obtaining Example-Based Explanations from Deep Neural Networks

Figure 2 for Obtaining Example-Based Explanations from Deep Neural Networks

Figure 3 for Obtaining Example-Based Explanations from Deep Neural Networks

Figure 4 for Obtaining Example-Based Explanations from Deep Neural Networks

Abstract:Most techniques for explainable machine learning focus on feature attribution, i.e., values are assigned to the features such that their sum equals the prediction. Example attribution is another form of explanation that assigns weights to the training examples, such that their scalar product with the labels equals the prediction. The latter may provide valuable complementary information to feature attribution, in particular in cases where the features are not easily interpretable. Current example-based explanation techniques have targeted a few model types only, such as k-nearest neighbors and random forests. In this work, a technique for obtaining example-based explanations from deep neural networks (EBE-DNN) is proposed. The basic idea is to use the deep neural network to obtain an embedding, which is employed by a k-nearest neighbor classifier to form a prediction; the example attribution can hence straightforwardly be derived from the latter. Results from an empirical investigation show that EBE-DNN can provide highly concentrated example attributions, i.e., the predictions can be explained with few training examples, without reducing accuracy compared to the original deep neural network. Another important finding from the empirical investigation is that the choice of layer to use for the embeddings may have a large impact on the resulting accuracy.

* To be published in the Symposium on Intelligent Data Analysis (IDA) 2025

Via

Access Paper or Ask Questions

Bitcoin Research with a Transaction Graph Dataset

Nov 15, 2024

Hugo Schnoering, Michalis Vazirgiannis

Figure 1 for Bitcoin Research with a Transaction Graph Dataset

Figure 2 for Bitcoin Research with a Transaction Graph Dataset

Figure 3 for Bitcoin Research with a Transaction Graph Dataset

Figure 4 for Bitcoin Research with a Transaction Graph Dataset

Abstract:Bitcoin, launched in 2008 by Satoshi Nakamoto, established a new digital economy where value can be stored and transferred in a fully decentralized manner - alleviating the need for a central authority. This paper introduces a large scale dataset in the form of a transactions graph representing transactions between Bitcoin users along with a set of tasks and baselines. The graph includes 252 million nodes and 785 million edges, covering a time span of nearly 13 years of and 670 million transactions. Each node and edge is timestamped. As for supervised tasks we provide two labeled sets i. a 33,000 nodes based on entity type and ii. nearly 100,000 Bitcoin addresses labeled with an entity name and an entity type. This is the largest publicly available data set of bitcoin transactions designed to facilitate advanced research and exploration in this domain, overcoming the limitations of existing datasets. Various graph neural network models are trained to predict node labels, establishing a baseline for future research. In addition, several use cases are presented to demonstrate the dataset's applicability beyond Bitcoin analysis. Finally, all data and source code is made publicly available to enable reproducibility of the results.

Via

Access Paper or Ask Questions

Gaussian Mixture Models Based Augmentation Enhances GNN Generalization

Nov 13, 2024

Yassine Abbahaddou, Fragkiskos D. Malliaros, Johannes F. Lutzeyer, Amine Mohamed Aboussalah, Michalis Vazirgiannis

Figure 1 for Gaussian Mixture Models Based Augmentation Enhances GNN Generalization

Figure 2 for Gaussian Mixture Models Based Augmentation Enhances GNN Generalization

Figure 3 for Gaussian Mixture Models Based Augmentation Enhances GNN Generalization

Figure 4 for Gaussian Mixture Models Based Augmentation Enhances GNN Generalization

Abstract:Graph Neural Networks (GNNs) have shown great promise in tasks like node and graph classification, but they often struggle to generalize, particularly to unseen or out-of-distribution (OOD) data. These challenges are exacerbated when training data is limited in size or diversity. To address these issues, we introduce a theoretical framework using Rademacher complexity to compute a regret bound on the generalization error and then characterize the effect of data augmentation. This framework informs the design of GMM-GDA, an efficient graph data augmentation (GDA) algorithm leveraging the capability of Gaussian Mixture Models (GMMs) to approximate any distribution. Our approach not only outperforms existing augmentation techniques in terms of generalization but also offers improved time complexity, making it highly suitable for real-world applications.

Via

Access Paper or Ask Questions

Post-Hoc Robustness Enhancement in Graph Neural Networks with Conditional Random Fields

Nov 08, 2024

Yassine Abbahaddou, Sofiane Ennadir, Johannes F. Lutzeyer, Fragkiskos D. Malliaros, Michalis Vazirgiannis

Figure 1 for Post-Hoc Robustness Enhancement in Graph Neural Networks with Conditional Random Fields

Figure 2 for Post-Hoc Robustness Enhancement in Graph Neural Networks with Conditional Random Fields

Figure 3 for Post-Hoc Robustness Enhancement in Graph Neural Networks with Conditional Random Fields

Figure 4 for Post-Hoc Robustness Enhancement in Graph Neural Networks with Conditional Random Fields

Abstract:Graph Neural Networks (GNNs), which are nowadays the benchmark approach in graph representation learning, have been shown to be vulnerable to adversarial attacks, raising concerns about their real-world applicability. While existing defense techniques primarily concentrate on the training phase of GNNs, involving adjustments to message passing architectures or pre-processing methods, there is a noticeable gap in methods focusing on increasing robustness during inference. In this context, this study introduces RobustCRF, a post-hoc approach aiming to enhance the robustness of GNNs at the inference stage. Our proposed method, founded on statistical relational learning using a Conditional Random Field, is model-agnostic and does not require prior knowledge about the underlying model architecture. We validate the efficacy of this approach across various models, leveraging benchmark node classification datasets.

Via

Access Paper or Ask Questions

Centrality Graph Shift Operators for Graph Neural Networks

Nov 07, 2024

Yassine Abbahaddou, Fragkiskos D. Malliaros, Johannes F. Lutzeyer, Michalis Vazirgiannis

Figure 1 for Centrality Graph Shift Operators for Graph Neural Networks

Figure 2 for Centrality Graph Shift Operators for Graph Neural Networks

Figure 3 for Centrality Graph Shift Operators for Graph Neural Networks

Figure 4 for Centrality Graph Shift Operators for Graph Neural Networks

Abstract:Graph Shift Operators (GSOs), such as the adjacency and graph Laplacian matrices, play a fundamental role in graph theory and graph representation learning. Traditional GSOs are typically constructed by normalizing the adjacency matrix by the degree matrix, a local centrality metric. In this work, we instead propose and study Centrality GSOs (CGSOs), which normalize adjacency matrices by global centrality metrics such as the PageRank, $k$-core or count of fixed length walks. We study spectral properties of the CGSOs, allowing us to get an understanding of their action on graph signals. We confirm this understanding by defining and running the spectral clustering algorithm based on different CGSOs on several synthetic and real-world datasets. We furthermore outline how our CGSO can act as the message passing operator in any Graph Neural Network and in particular demonstrate strong performance of a variant of the Graph Convolutional Network and Graph Attention Network using our CGSOs on several real-world benchmark datasets.

Via

Access Paper or Ask Questions

Graph Neural Networks on Discriminative Graphs of Words

Oct 27, 2024

Yassine Abbahaddou, Johannes F. Lutzeyer, Michalis Vazirgiannis

Figure 1 for Graph Neural Networks on Discriminative Graphs of Words

Figure 2 for Graph Neural Networks on Discriminative Graphs of Words

Figure 3 for Graph Neural Networks on Discriminative Graphs of Words

Figure 4 for Graph Neural Networks on Discriminative Graphs of Words

Abstract:In light of the recent success of Graph Neural Networks (GNNs) and their ability to perform inference on complex data structures, many studies apply GNNs to the task of text classification. In most previous methods, a heterogeneous graph, containing both word and document nodes, is constructed using the entire corpus and a GNN is used to classify document nodes. In this work, we explore a new Discriminative Graph of Words Graph Neural Network (DGoW-GNN) approach encapsulating both a novel discriminative graph construction and model to classify text. In our graph construction, containing only word nodes and no document nodes, we split the training corpus into disconnected subgraphs according to their labels and weight edges by the pointwise mutual information of the represented words. Our graph construction, for which we provide theoretical motivation, allows us to reformulate the task of text classification as the task of walk classification. We also propose a new model for the graph-based classification of text, which combines a GNN and a sequence model. We evaluate our approach on seven benchmark datasets and find that it is outperformed by several state-of-the-art baseline models. We analyse reasons for this performance difference and hypothesise under which conditions it is likely to change.

Via

Access Paper or Ask Questions

Graph Linearization Methods for Reasoning on Graphs with Large Language Models

Oct 25, 2024

Christos Xypolopoulos, Guokan Shang, Xiao Fei, Giannis Nikolentzos, Hadi Abdine, Iakovos Evdaimon, Michail Chatzianastasis, Giorgos Stamou, Michalis Vazirgiannis

Figure 1 for Graph Linearization Methods for Reasoning on Graphs with Large Language Models

Figure 2 for Graph Linearization Methods for Reasoning on Graphs with Large Language Models

Figure 3 for Graph Linearization Methods for Reasoning on Graphs with Large Language Models

Figure 4 for Graph Linearization Methods for Reasoning on Graphs with Large Language Models

Abstract:Large language models have evolved to process multiple modalities beyond text, such as images and audio, which motivates us to explore how to effectively leverage them for graph machine learning tasks. The key question, therefore, is how to transform graphs into linear sequences of tokens, a process we term graph linearization, so that LLMs can handle graphs naturally. We consider that graphs should be linearized meaningfully to reflect certain properties of natural language text, such as local dependency and global alignment, in order to ease contemporary LLMs, trained on trillions of textual tokens, better understand graphs. To achieve this, we developed several graph linearization methods based on graph centrality, degeneracy, and node relabeling schemes. We then investigated their effect on LLM performance in graph reasoning tasks. Experimental results on synthetic graphs demonstrate the effectiveness of our methods compared to random linearization baselines. Our work introduces novel graph representations suitable for LLMs, contributing to the potential integration of graph machine learning with the trend of multi-modal processing using a unified transformer model.

Via

Access Paper or Ask Questions

Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks ?

Oct 17, 2024

Virgile Rennard, Christos Xypolopoulos, Michalis Vazirgiannis

Figure 1 for Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks ?

Figure 2 for Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks ?

Figure 3 for Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks ?

Figure 4 for Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks ?

Abstract:Large language models (LLMs) inherit biases from their training data and alignment processes, influencing their responses in subtle ways. While many studies have examined these biases, little work has explored their robustness during interactions. In this paper, we introduce a novel approach where two instances of an LLM engage in self-debate, arguing opposing viewpoints to persuade a neutral version of the model. Through this, we evaluate how firmly biases hold and whether models are susceptible to reinforcing misinformation or shifting to harmful viewpoints. Our experiments span multiple LLMs of varying sizes, origins, and languages, providing deeper insights into bias persistence and flexibility across linguistic and cultural contexts.

Via

Access Paper or Ask Questions

Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect

Sep 26, 2024

Guokan Shang, Hadi Abdine, Yousef Khoubrane, Amr Mohamed, Yassine Abbahaddou, Sofiane Ennadir, Imane Momayiz, Xuguang Ren, Eric Moulines, Preslav Nakov(+2 more)

Abstract:We introduce Atlas-Chat, the first-ever collection of large language models specifically developed for dialectal Arabic. Focusing on Moroccan Arabic, also known as Darija, we construct our instruction dataset by consolidating existing Darija language resources, creating novel datasets both manually and synthetically, and translating English instructions with stringent quality control. Atlas-Chat-9B and 2B models, fine-tuned on the dataset, exhibit superior ability in following Darija instructions and performing standard NLP tasks. Notably, our models outperform both state-of-the-art and Arabic-specialized LLMs like LLaMa, Jais, and AceGPT, e.g., achieving a 13% performance boost over a larger 13B model on DarijaMMLU, in our newly introduced evaluation suite for Darija covering both discriminative and generative tasks. Furthermore, we perform an experimental analysis of various fine-tuning strategies and base model choices to determine optimal configurations. All our resources are publicly accessible, and we believe our work offers comprehensive design methodologies of instruction-tuning for low-resource language variants, which are often neglected in favor of data-rich languages by contemporary LLMs.

Via

Access Paper or Ask Questions