Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Lei Li

Carnegie Mellon University

INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback

May 23, 2023

Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Yang Wang, Lei Li

Figure 1 for INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback

Figure 2 for INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback

Figure 3 for INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback

Figure 4 for INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback

Abstract:The field of automatic evaluation of text generation made tremendous progress in the last few years. In particular, since the advent of neural metrics, like COMET, BLEURT, and SEScore2, the newest generation of metrics show a high correlation with human judgment. Unfortunately, quality scores generated with neural metrics are not interpretable, and it is unclear which part of the generation output is criticized by the metrics. To address this limitation, we present INSTRUCTSCORE, an open-source, explainable evaluation metric for text generation. By harnessing both explicit human instruction and the implicit knowledge of GPT4, we fine-tune a LLAMA model to create an evaluative metric that can produce a diagnostic report aligned with human judgment. We evaluate INSTRUCTSCORE on the WMT22 Zh-En translation task, where our 7B model surpasses other LLM-based baselines, including those based on 175B GPT3. Impressively, our INSTRUCTSCORE, even without direct supervision from human-rated data, achieves performance levels on par with state-of-the-art metrics like COMET22, which was fine-tuned on human ratings.

* Work in progress

Via

Access Paper or Ask Questions

Learn from Mistakes through Cooperative Interaction with Study Assistant

May 23, 2023

Danqing Wang, Lei Li

Figure 1 for Learn from Mistakes through Cooperative Interaction with Study Assistant

Figure 2 for Learn from Mistakes through Cooperative Interaction with Study Assistant

Figure 3 for Learn from Mistakes through Cooperative Interaction with Study Assistant

Figure 4 for Learn from Mistakes through Cooperative Interaction with Study Assistant

Abstract:Large language models have demonstrated their ability to self-reflect and refine their generation, which can further improve their performance. However, this feedback mechanism faces challenges such as no guarantee of correctness and the lack of global insight into the model's weaknesses. In this paper, we propose a novel framework, Study Assistant for Large Language Model (SALAM), to aid LLMs in the reflection and refinement process. Motivated by the human study assistant, this framework grades previous responses with the ground truth and collects mistakes in the training phase. During inference, it identifies common misunderstandings based on the mistake collections and provides guidelines for the model to help the model avoid similar mistakes during inference. SALAM is a model-agnostic framework, focusing on providing general feedback and can adapt to any base model. Our evaluation of SALAM on two challenging benchmarks demonstrated a significant improvement over various baselines.

Via

Access Paper or Ask Questions

Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

May 23, 2023

Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun

Figure 1 for Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Figure 2 for Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Figure 3 for Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Figure 4 for Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Abstract:In-context learning (ICL) emerges as a promising capability of large language models (LLMs) by providing them with demonstration examples to perform diverse tasks. However, the underlying mechanism of how LLMs learn from the provided context remains under-explored. In this paper, we investigate the working mechanism of ICL through an information flow lens. Our findings reveal that label words in the demonstration examples function as anchors: (1) semantic information aggregates into label word representations during the shallow computation layers' processing; (2) the consolidated information in label words serves as a reference for LLMs' final predictions. Based on these insights, we introduce an anchor re-weighting method to improve ICL performance, a demonstration compression technique to expedite inference, and an analysis framework for diagnosing ICL errors in GPT2-XL. The promising applications of our findings again validate the uncovered ICL working mechanism and pave the way for future studies.

Via

Access Paper or Ask Questions

Can We Edit Factual Knowledge by In-Context Learning?

May 22, 2023

Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, Baobao Chang

Figure 1 for Can We Edit Factual Knowledge by In-Context Learning?

Figure 2 for Can We Edit Factual Knowledge by In-Context Learning?

Figure 3 for Can We Edit Factual Knowledge by In-Context Learning?

Figure 4 for Can We Edit Factual Knowledge by In-Context Learning?

Abstract:Previous studies have shown that large language models (LLMs) like GPTs store massive factual knowledge in their parameters. However, the stored knowledge could be false or out-dated. Traditional knowledge editing methods refine LLMs via fine-tuning on texts containing specific knowledge. However, with the increasing scales of LLMs, these gradient-based approaches bring large computation costs. The trend of model-as-a-service also makes it impossible to modify knowledge in black-box LMs. Inspired by in-context learning (ICL), a new paradigm based on demonstration contexts without parameter updating, we explore whether ICL can edit factual knowledge. To answer this question, we give a comprehensive empirical study of ICL strategies. Experiments show that in-context knowledge editing (IKE), without any gradient and parameter updating, achieves a competitive success rate compared to gradient-based methods on GPT-J (6B) but with much fewer side effects, including less over-editing on similar but unrelated facts and less knowledge forgetting on previously stored knowledge. We also apply the method to larger LMs with tens or hundreds of parameters like OPT-175B, which shows the scalability of our method. The code is available at https://github.com/Zce1112zslx/IKE.

Via

Access Paper or Ask Questions

Extrapolating Multilingual Understanding Models as Multilingual Generators

May 22, 2023

Bohong Wu, Fei Yuan, Hai Zhao, Lei Li, Jingjing Xu

Figure 1 for Extrapolating Multilingual Understanding Models as Multilingual Generators

Figure 2 for Extrapolating Multilingual Understanding Models as Multilingual Generators

Figure 3 for Extrapolating Multilingual Understanding Models as Multilingual Generators

Figure 4 for Extrapolating Multilingual Understanding Models as Multilingual Generators

Abstract:Multilingual understanding models (or encoder-based), pre-trained via masked language modeling, have achieved promising results on many language understanding tasks (e.g., mBERT). However, these non-autoregressive (NAR) models still struggle to generate high-quality texts compared with autoregressive (AR) models. Considering that encoder-based models have the advantage of efficient generation and self-correction abilities, this paper explores methods to empower multilingual understanding models the generation abilities to get a unified model. Specifically, we start from a multilingual encoder (XLM-R) and propose a \textbf{S}emantic-\textbf{G}uided \textbf{A}lignment-then-Denoising (SGA) approach to adapt an encoder to a multilingual generator with a small number of new parameters. Experiments show that the proposed approach is an effective adaption method, outperforming widely-used initialization-based methods with gains of 9.4 BLEU on machine translation, 8.1 Rouge-L on question generation, and 5.5 METEOR on story generation on XLM-R$_{large}$. On the other hand, we observe that XLM-R is still inferior to mBART in supervised settings despite better results on zero-shot settings, indicating that more exploration is required to make understanding models strong generators.

Via

Access Paper or Ask Questions

Communication Efficient Federated Learning for Multilingual Neural Machine Translation with Adapter

May 21, 2023

Yi Liu, Xiaohan Bi, Lei Li, Sishuo Chen, Wenkai Yang, Xu Sun

Figure 1 for Communication Efficient Federated Learning for Multilingual Neural Machine Translation with Adapter

Figure 2 for Communication Efficient Federated Learning for Multilingual Neural Machine Translation with Adapter

Figure 3 for Communication Efficient Federated Learning for Multilingual Neural Machine Translation with Adapter

Figure 4 for Communication Efficient Federated Learning for Multilingual Neural Machine Translation with Adapter

Abstract:Federated Multilingual Neural Machine Translation (Fed-MNMT) has emerged as a promising paradigm for institutions with limited language resources. This approach allows multiple institutions to act as clients and train a unified model through model synchronization, rather than collecting sensitive data for centralized training. This significantly reduces the cost of corpus collection and preserves data privacy. However, as pre-trained language models (PLMs) continue to increase in size, the communication cost for transmitting parameters during synchronization has become a training speed bottleneck. In this paper, we propose a communication-efficient Fed-MNMT framework that addresses this issue by keeping PLMs frozen and only transferring lightweight adapter modules between clients. Since different language pairs exhibit substantial discrepancies in data distributions, adapter parameters of clients may conflict with each other. To tackle this, we explore various clustering strategies to group parameters for integration and mitigate the negative effects of conflicting parameters. Experimental results demonstrate that our framework reduces communication cost by over 98% while achieving similar or even better performance compared to competitive baselines. Further analysis reveals that clustering strategies effectively solve the problem of linguistic discrepancy and pruning adapter modules further improves communication efficiency.

* Findings of ACL 2023

Via

Access Paper or Ask Questions

Statistical Knowledge Assessment for Generative Language Models

May 17, 2023

Qingxiu Dong, Jingjing Xu, Lingpeng Kong, Zhifang Sui, Lei Li

Figure 1 for Statistical Knowledge Assessment for Generative Language Models

Figure 2 for Statistical Knowledge Assessment for Generative Language Models

Figure 3 for Statistical Knowledge Assessment for Generative Language Models

Figure 4 for Statistical Knowledge Assessment for Generative Language Models

Abstract:Generative Language Models (GLMs) have demonstrated capabilities to store factual knowledge and answer queries efficiently. Given varying prompts, does a GLM consistently generate factually correct answers? In this paper, we introduce a statistical knowledge assessment framework guided by latent variables and the KaRR metric, which quantifies a model's knowledge by computing its continuous probability across diverse text forms. We conduct a comprehensive comparison of knowledge across 14 GLMs using our framework, including LLaMA, Alpaca, OPT, and others. Our statistical knowledge assessment encompasses 600 relation types and exhibits a strong correlation (0.43 Kendall's $\tau$) with human evaluation. Our findings reveal that the knowledge in GLMs with the same backbone architecture adheres to the scaling law, and that tuning on instruction-following data may compromise the model's ability to generate factually correct text consistently.

Via

Access Paper or Ask Questions

Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

May 13, 2023

Jiangjie Chen, Wei Shi, Ziquan Fu, Sijie Cheng, Lei Li, Yanghua Xiao

Figure 1 for Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

Figure 2 for Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

Figure 3 for Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

Figure 4 for Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

Abstract:Large language models (LLMs) have been widely studied for their ability to store and utilize positive knowledge. However, negative knowledge, such as "lions don't live in the ocean", is also ubiquitous in the world but rarely mentioned explicitly in the text. What do LLMs know about negative knowledge? This work examines the ability of LLMs to negative commonsense knowledge. We design a constrained keywords-to-sentence generation task (CG) and a Boolean question-answering task (QA) to probe LLMs. Our experiments reveal that LLMs frequently fail to generate valid sentences grounded in negative commonsense knowledge, yet they can correctly answer polar yes-or-no questions. We term this phenomenon the belief conflict of LLMs. Our further analysis shows that statistical shortcuts and negation reporting bias from language modeling pre-training cause this conflict.

* Accepted to ACL 2023

Via

Access Paper or Ask Questions

HIORE: Leveraging High-order Interactions for Unified Entity Relation Extraction

May 07, 2023

Yijun Wang, Changzhi Sun, Yuanbin Wu, Lei Li, Junchi Yan, Hao Zhou

Abstract:Entity relation extraction consists of two sub-tasks: entity recognition and relation extraction. Existing methods either tackle these two tasks separately or unify them with word-by-word interactions. In this paper, we propose HIORE, a new method for unified entity relation extraction. The key insight is to leverage the high-order interactions, i.e., the complex association among word pairs, which contains richer information than the first-order word-by-word interactions. For this purpose, we first devise a W-shape DNN (WNet) to capture coarse-level high-order connections. Then, we build a heuristic high-order graph and further calibrate the representations with a graph neural network (GNN). Experiments on three benchmarks (ACE04, ACE05, SciERC) show that HIORE achieves the state-of-the-art performance on relation extraction and an improvement of 1.1~1.8 F1 points over the prior best unified model.

* 10 pages

Via

Access Paper or Ask Questions

Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis

May 02, 2023

Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, Lei Li

Figure 1 for Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis

Figure 2 for Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis

Figure 3 for Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis

Figure 4 for Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis

Abstract:Large language models (LLMs) have demonstrated remarkable potential in handling multilingual machine translation (MMT). In this paper, we systematically investigate the advantages and challenges of LLMs for MMT by answering two questions: 1) How well do LLMs perform in translating a massive number of languages? 2) Which factors affect LLMs' performance in translation? We evaluate popular LLMs, including XGLM, OPT, BLOOMZ, and ChatGPT, on 102 languages. Our empirical results show that even the best model ChatGPT still lags behind the supervised baseline NLLB in 83.33% of translation directions. Through further analysis, we discover that LLMs exhibit new working patterns when used for MMT. First, prompt semantics can surprisingly be ignored when given in-context exemplars, where LLMs still show strong performance even with unreasonable prompts. Second, cross-lingual exemplars can provide better task instruction for low-resource translation than exemplars in the same language pairs. Third, we observe the overestimated performance of BLOOMZ on dataset Flores-101, indicating the potential risk when using public datasets for evaluation.

Via

Access Paper or Ask Questions