Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Lin Wang

Rethinking Knowledge in Distillation: An In-context Sample Retrieval Perspective

Jan 13, 2025

Jinjing Zhu, Songze Li, Lin Wang

Figure 1 for Rethinking Knowledge in Distillation: An In-context Sample Retrieval Perspective

Figure 2 for Rethinking Knowledge in Distillation: An In-context Sample Retrieval Perspective

Figure 3 for Rethinking Knowledge in Distillation: An In-context Sample Retrieval Perspective

Figure 4 for Rethinking Knowledge in Distillation: An In-context Sample Retrieval Perspective

Abstract:Conventional knowledge distillation (KD) approaches are designed for the student model to predict similar output as the teacher model for each sample. Unfortunately, the relationship across samples with same class is often neglected. In this paper, we explore to redefine the knowledge in distillation, capturing the relationship between each sample and its corresponding in-context samples (a group of similar samples with the same or different classes), and perform KD from an in-context sample retrieval perspective. As KD is a type of learned label smoothing regularization (LSR), we first conduct a theoretical analysis showing that the teacher's knowledge from the in-context samples is a crucial contributor to regularize the student training with the corresponding samples. Buttressed by the analysis, we propose a novel in-context knowledge distillation (IC-KD) framework that shows its superiority across diverse KD paradigms (offline, online, and teacher-free KD). Firstly, we construct a feature memory bank from the teacher model and retrieve in-context samples for each corresponding sample through retrieval-based learning. We then introduce Positive In-Context Distillation (PICD) to reduce the discrepancy between a sample from the student and the aggregated in-context samples with the same class from the teacher in the logit space. Moreover, Negative In-Context Distillation (NICD) is introduced to separate a sample from the student and the in-context samples with different classes from the teacher in the logit space. Extensive experiments demonstrate that IC-KD is effective across various types of KD, and consistently achieves state-of-the-art performance on CIFAR-100 and ImageNet datasets.

Via

Access Paper or Ask Questions

Efficient Graph Condensation via Gaussian Process

Jan 05, 2025

Lin Wang, Qing Li

Figure 1 for Efficient Graph Condensation via Gaussian Process

Figure 2 for Efficient Graph Condensation via Gaussian Process

Figure 3 for Efficient Graph Condensation via Gaussian Process

Figure 4 for Efficient Graph Condensation via Gaussian Process

Abstract:Graph condensation reduces the size of large graphs while preserving performance, addressing the scalability challenges of Graph Neural Networks caused by computational inefficiencies on large datasets. Existing methods often rely on bi-level optimization, requiring extensive GNN training and limiting their scalability. To address these issues, this paper proposes Graph Condensation via Gaussian Process (GCGP), a novel and computationally efficient approach to graph condensation. GCGP utilizes a Gaussian Process (GP), with the condensed graph serving as observations, to estimate the posterior distribution of predictions. This approach eliminates the need for the iterative and resource-intensive training typically required by GNNs. To enhance the capability of the GCGP in capturing dependencies between function values, we derive a specialized covariance function that incorporates structural information. This covariance function broadens the receptive field of input nodes by local neighborhood aggregation, thereby facilitating the representation of intricate dependencies within the nodes. To address the challenge of optimizing binary structural information in condensed graphs, Concrete random variables are utilized to approximate the binary adjacency matrix in a continuous counterpart. This relaxation process allows the adjacency matrix to be represented in a differentiable form, enabling the application of gradient-based optimization techniques to discrete graph structures. Experimental results show that the proposed GCGP method efficiently condenses large-scale graph data while preserving predictive performance, addressing the scalability and efficiency challenges. The implementation of our method is publicly available at https://github.com/WANGLin0126/GCGP.

Via

Access Paper or Ask Questions

MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection

Dec 22, 2024

Xu Zheng, Yuanhuiyi Lyu, Lutao Jiang, Jiazhou Zhou, Lin Wang, Xuming Hu

Abstract:In this paper, we address the challenging modality-agnostic semantic segmentation (MaSS), aiming at centering the value of every modality at every feature granularity. Training with all available visual modalities and effectively fusing an arbitrary combination of them is essential for robust multi-modal fusion in semantic segmentation, especially in real-world scenarios, yet remains less explored to date. Existing approaches often place RGB at the center, treating other modalities as secondary, resulting in an asymmetric architecture. However, RGB alone can be limiting in scenarios like nighttime, where modalities such as event data excel. Therefore, a resilient fusion model must dynamically adapt to each modality's strengths while compensating for weaker inputs.To this end, we introduce the MAGIC++ framework, which comprises two key plug-and-play modules for effective multi-modal fusion and hierarchical modality selection that can be equipped with various backbone models. Firstly, we introduce a multi-modal interaction module to efficiently process features from the input multi-modal batches and extract complementary scene information with channel-wise and spatial-wise guidance. On top, a unified multi-scale arbitrary-modal selection module is proposed to utilize the aggregated features as the benchmark to rank the multi-modal features based on the similarity scores at hierarchical feature spaces. This way, our method can eliminate the dependence on RGB modality at every feature granularity and better overcome sensor failures and environmental noises while ensuring the segmentation performance. Under the common multi-modal setting, our method achieves state-of-the-art performance on both real-world and synthetic benchmarks. Moreover, our method is superior in the novel modality-agnostic setting, where it outperforms prior arts by a large margin.

Via

Access Paper or Ask Questions

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Dec 20, 2024

Xingchen Song, Chengdong Liang, Binbin Zhang, Pengshen Zhang, ZiYu Wang, Youcheng Ma, Menglong Xu, Lin Wang, Di Wu, Fuping Pan(+2 more)

Figure 1 for TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Figure 2 for TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Figure 3 for TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Figure 4 for TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Abstract:Large Automatic Speech Recognition (ASR) models demand a vast number of parameters, copious amounts of data, and significant computational resources during the training process. However, such models can merely be deployed on high-compute cloud platforms and are only capable of performing speech recognition tasks. This leads to high costs and restricted capabilities. In this report, we initially propose the elastic mixture of the expert (eMoE) model. This model can be trained just once and then be elastically scaled in accordance with deployment requirements. Secondly, we devise an unsupervised data creation and validation procedure and gather millions of hours of audio data from diverse domains for training. Using these two techniques, our system achieves elastic deployment capabilities while reducing the Character Error Rate (CER) on the SpeechIO testsets from 4.98\% to 2.45\%. Thirdly, our model is not only competent in Mandarin speech recognition but also proficient in multilingual, multi-dialect, emotion, gender, and sound event perception. We refer to this as Automatic Speech Perception (ASP), and the perception results are presented in the experimental section.

* Technical Report

Via

Access Paper or Ask Questions

MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules

Dec 18, 2024

Kejie Chen, Lin Wang, Qinghai Zhang, Renjun Xu

Figure 1 for MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules

Figure 2 for MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules

Figure 3 for MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules

Figure 4 for MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules

Abstract:Recent studies have highlighted the limitations of large language models in mathematical reasoning, particularly their inability to capture the underlying logic. Inspired by meta-learning, we propose that models should acquire not only task-specific knowledge but also transferable problem-solving skills. We introduce MetaRuleGPT, a novel Transformer-based architecture that performs precise numerical calculations and complex logical operations by learning and combining different rules. In contrast with traditional training sets, which are heavily composed of massive raw instance data, MetaRuleGPT is pre-trained on much less abstract datasets containing basic, compound, and iterative rules for mathematical reasoning. Extensive experimental results demonstrate MetaRuleGPT can mimic human's rule-following capabilities, break down complexity, and iteratively derive accurate results for complex mathematical problems. These findings prove the potential of rule learning to enhance the numerical reasoning abilities of language models.

* 8 pages, 6 figures

Via

Access Paper or Ask Questions

Enhanced 3D Generation by 2D Editing

Dec 08, 2024

Haoran Li, Yuli Tian, Yong Liao, Lin Wang, Yuyang Wang, Peng Yuan Zhou

Figure 1 for Enhanced 3D Generation by 2D Editing

Figure 2 for Enhanced 3D Generation by 2D Editing

Figure 3 for Enhanced 3D Generation by 2D Editing

Figure 4 for Enhanced 3D Generation by 2D Editing

Abstract:Distilling 3D representations from pretrained 2D diffusion models is essential for 3D creative applications across gaming, film, and interior design. Current SDS-based methods are hindered by inefficient information distillation from diffusion models, which prevents the creation of photorealistic 3D contents. Our research reevaluates the SDS approach by analyzing its fundamental nature as a basic image editing process that commonly results in over-saturation, over-smoothing and lack of rich content due to the poor-quality single-step denoising. To address these limitations, we propose GE3D (3D Generation by Editing). Each iteration of GE3D utilizes a 2D editing framework that combines a noising trajectory to preserve the information of the input image, alongside a text-guided denoising trajectory. We optimize the process by aligning the latents across both trajectories. This approach fully exploits pretrained diffusion models to distill multi-granularity information through multiple denoising steps, resulting in photorealistic 3D outputs. Both theoretical and experimental results confirm the effectiveness of our approach, which not only advances 3D generation technology but also establishes a novel connection between 3D generation and 2D editing. This could potentially inspire further research in the field. Code and demos are released at https://jahnsonblack.github.io/GE3D/.

Via

Access Paper or Ask Questions

Recent Trends in Linear Text Segmentation: a Survey

Nov 25, 2024

Iacopo Ghinassi, Lin Wang, Chris Newell, Matthew Purver

Figure 1 for Recent Trends in Linear Text Segmentation: a Survey

Figure 2 for Recent Trends in Linear Text Segmentation: a Survey

Figure 3 for Recent Trends in Linear Text Segmentation: a Survey

Figure 4 for Recent Trends in Linear Text Segmentation: a Survey

Abstract:Linear Text Segmentation is the task of automatically tagging text documents with topic shifts, i.e. the places in the text where the topics change. A well-established area of research in Natural Language Processing, drawing from well-understood concepts in linguistic and computational linguistic research, the field has recently seen a lot of interest as a result of the surge of text, video, and audio available on the web, which in turn require ways of summarising and categorizing the mole of content for which linear text segmentation is a fundamental step. In this survey, we provide an extensive overview of current advances in linear text segmentation, describing the state of the art in terms of resources and approaches for the task. Finally, we highlight the limitations of available resources and of the task itself, while indicating ways forward based on the most recent literature and under-explored research directions.

* Findings of the Association for Computational Linguistics: EMNLP 2024

Via

Access Paper or Ask Questions

DiffRoad: Realistic and Diverse Road Scenario Generation for Autonomous Vehicle Testing

Nov 14, 2024

Junjie Zhou, Lin Wang, Qiang Meng, Xiaofan Wang

Figure 1 for DiffRoad: Realistic and Diverse Road Scenario Generation for Autonomous Vehicle Testing

Figure 2 for DiffRoad: Realistic and Diverse Road Scenario Generation for Autonomous Vehicle Testing

Figure 3 for DiffRoad: Realistic and Diverse Road Scenario Generation for Autonomous Vehicle Testing

Figure 4 for DiffRoad: Realistic and Diverse Road Scenario Generation for Autonomous Vehicle Testing

Abstract:Generating realistic and diverse road scenarios is essential for autonomous vehicle testing and validation. Nevertheless, owing to the complexity and variability of real-world road environments, creating authentic and varied scenarios for intelligent driving testing is challenging. In this paper, we propose DiffRoad, a novel diffusion model designed to produce controllable and high-fidelity 3D road scenarios. DiffRoad leverages the generative capabilities of diffusion models to synthesize road layouts from white noise through an inverse denoising process, preserving real-world spatial features. To enhance the quality of generated scenarios, we design the Road-UNet architecture, optimizing the balance between backbone and skip connections for high-realism scenario generation. Furthermore, we introduce a road scenario evaluation module that screens adequate and reasonable scenarios for intelligent driving testing using two critical metrics: road continuity and road reasonableness. Experimental results on multiple real-world datasets demonstrate DiffRoad's ability to generate realistic and smooth road structures while maintaining the original distribution. Additionally, the generated scenarios can be fully automated into the OpenDRIVE format, facilitating generalized autonomous vehicle simulation testing. DiffRoad provides a rich and diverse scenario library for large-scale autonomous vehicle testing and offers valuable insights for future infrastructure designs that are better suited for autonomous vehicles.

* 14 pages, 9 figures

Via

Access Paper or Ask Questions

Generative Emotion Cause Explanation in Multimodal Conversations

Nov 01, 2024

Lin Wang, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang

Figure 1 for Generative Emotion Cause Explanation in Multimodal Conversations

Figure 2 for Generative Emotion Cause Explanation in Multimodal Conversations

Figure 3 for Generative Emotion Cause Explanation in Multimodal Conversations

Figure 4 for Generative Emotion Cause Explanation in Multimodal Conversations

Abstract:Multimodal conversation, a crucial form of human communication, carries rich emotional content, making the exploration of the causes of emotions within it a research endeavor of significant importance. However, existing research on the causes of emotions typically uses clause selection methods to locate the reason utterance, without providing a detailed explanation of the emotional causes. In this paper, we propose a new task, \textbf{M}ultimodal \textbf{C}onversation \textbf{E}motion \textbf{C}ause \textbf{E}xplanation (MCECE), aiming to generate a detailed explanation of the emotional cause to the target utterance within a multimodal conversation scenario. Building upon the MELD dataset, we develop a new dataset (ECEM) that integrates video clips with detailed explanations of character emotions, facilitating an in-depth examination of the causal factors behind emotional expressions in multimodal conversations.A novel approach, FAME-Net, is further proposed, that harnesses the power of Large Language Models (LLMs) to analyze visual data and accurately interpret the emotions conveyed through facial expressions in videos. By exploiting the contagion effect of facial emotions, FAME-Net effectively captures the emotional causes of individuals engaged in conversations. Our experimental results on the newly constructed dataset show that FAME-Net significantly outperforms several excellent large language model baselines. Code and dataset are available at \url{https://github.com/3222345200/ECEMdataset.git}

Via

Access Paper or Ask Questions

FedMABA: Towards Fair Federated Learning through Multi-Armed Bandits Allocation

Oct 26, 2024

Zhichao Wang, Lin Wang, Yongxin Guo, Ying-Jun Angela Zhang, Xiaoying Tang

Figure 1 for FedMABA: Towards Fair Federated Learning through Multi-Armed Bandits Allocation

Figure 2 for FedMABA: Towards Fair Federated Learning through Multi-Armed Bandits Allocation

Figure 3 for FedMABA: Towards Fair Federated Learning through Multi-Armed Bandits Allocation

Figure 4 for FedMABA: Towards Fair Federated Learning through Multi-Armed Bandits Allocation

Abstract:The increasing concern for data privacy has driven the rapid development of federated learning (FL), a privacy-preserving collaborative paradigm. However, the statistical heterogeneity among clients in FL results in inconsistent performance of the server model across various clients. Server model may show favoritism towards certain clients while performing poorly for others, heightening the challenge of fairness. In this paper, we reconsider the inconsistency in client performance distribution and introduce the concept of adversarial multi-armed bandit to optimize the proposed objective with explicit constraints on performance disparities. Practically, we propose a novel multi-armed bandit-based allocation FL algorithm (FedMABA) to mitigate performance unfairness among diverse clients with different data distributions. Extensive experiments, in different Non-I.I.D. scenarios, demonstrate the exceptional performance of FedMABA in enhancing fairness.

Via

Access Paper or Ask Questions