Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Bo Han

Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection

Nov 05, 2024

Geng Yu, Jianing Zhu, Jiangchao Yao, Bo Han

Figure 1 for Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection

Figure 2 for Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection

Figure 3 for Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection

Figure 4 for Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection

Abstract:Out-of-distribution (OOD) detection is crucial for deploying reliable machine learning models in open-world applications. Recent advances in CLIP-based OOD detection have shown promising results via regularizing prompt tuning with OOD features extracted from ID data. However, the irrelevant context mined from ID data can be spurious due to the inaccurate foreground-background decomposition, thus limiting the OOD detection performance. In this work, we propose a novel framework, namely, Self-Calibrated Tuning (SCT), to mitigate this problem for effective OOD detection with only the given few-shot ID data. Specifically, SCT introduces modulating factors respectively on the two components of the original learning objective. It adaptively directs the optimization process between the two tasks during training on data with different prediction uncertainty to calibrate the influence of OOD regularization, which is compatible with many prompt tuning based OOD detection methods. Extensive experiments and analyses have been conducted to characterize and demonstrate the effectiveness of the proposed SCT. The code is publicly available.

* accepted by NeurIPS 2024

Via

Access Paper or Ask Questions

Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?

Oct 31, 2024

Zhanke Zhou, Rong Tao, Jianing Zhu, Yiwen Luo, Zengmao Wang, Bo Han

Figure 1 for Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?

Figure 2 for Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?

Figure 3 for Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?

Figure 4 for Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?

Abstract:This paper investigates an under-explored challenge in large language models (LLMs): chain-of-thought prompting with noisy rationales, which include irrelevant or inaccurate reasoning thoughts within examples used for in-context learning. We construct NoRa dataset that is tailored to evaluate the robustness of reasoning in the presence of noisy rationales. Our findings on NoRa dataset reveal a prevalent vulnerability to such noise among current LLMs, with existing robust methods like self-correction and self-consistency showing limited efficacy. Notably, compared to prompting with clean rationales, base LLM drops by 1.4%-19.8% in accuracy with irrelevant thoughts and more drastically by 2.2%-40.4% with inaccurate thoughts. Addressing this challenge necessitates external supervision that should be accessible in practice. Here, we propose the method of contrastive denoising with noisy chain-of-thought (CD-CoT). It enhances LLMs' denoising-reasoning capabilities by contrasting noisy rationales with only one clean rationale, which can be the minimal requirement for denoising-purpose prompting. This method follows a principle of exploration and exploitation: (1) rephrasing and selecting rationales in the input space to achieve explicit denoising and (2) exploring diverse reasoning paths and voting on answers in the output space. Empirically, CD-CoT demonstrates an average improvement of 17.8% in accuracy over the base model and shows significantly stronger denoising capabilities than baseline methods. The source code is publicly available at: https://github.com/tmlr-group/NoisyRationales.

* Accepted by NeurIPS 2024

Via

Access Paper or Ask Questions

FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion

Oct 27, 2024

Zhenheng Tang, Yonggang Zhang, Peijie Dong, Yiu-ming Cheung, Amelie Chi Zhou, Bo Han, Xiaowen Chu

Figure 1 for FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion

Figure 2 for FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion

Figure 3 for FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion

Figure 4 for FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion

Abstract:One-shot Federated Learning (OFL) significantly reduces communication costs in FL by aggregating trained models only once. However, the performance of advanced OFL methods is far behind the normal FL. In this work, we provide a causal view to find that this performance drop of OFL methods comes from the isolation problem, which means that local isolatedly trained models in OFL may easily fit to spurious correlations due to the data heterogeneity. From the causal perspective, we observe that the spurious fitting can be alleviated by augmenting intermediate features from other clients. Built upon our observation, we propose a novel learning approach to endow OFL with superb performance and low communication and storage costs, termed as FuseFL. Specifically, FuseFL decomposes neural networks into several blocks, and progressively trains and fuses each block following a bottom-up manner for feature augmentation, introducing no additional communication costs. Comprehensive experiments demonstrate that FuseFL outperforms existing OFL and ensemble FL by a significant margin. We conduct comprehensive experiments to show that FuseFL supports high scalability of clients, heterogeneous model training, and low memory costs. Our work is the first attempt using causality to analyze and alleviate data heterogeneity of OFL.

Via

Access Paper or Ask Questions

Irregular Tensor Low-Rank Representation for Hyperspectral Image Representation

Oct 24, 2024

Bo Han, Yuheng Jia, Hui Liu, Junhui Hou

Figure 1 for Irregular Tensor Low-Rank Representation for Hyperspectral Image Representation

Figure 2 for Irregular Tensor Low-Rank Representation for Hyperspectral Image Representation

Figure 3 for Irregular Tensor Low-Rank Representation for Hyperspectral Image Representation

Figure 4 for Irregular Tensor Low-Rank Representation for Hyperspectral Image Representation

Abstract:Spectral variation is a common problem for hyperspectral image (HSI) representation. Low-rank tensor representation is an important approach to alleviate spectral variations. However, the spatial distribution of the HSI is always irregular, while the previous tensor low-rank representation methods can only be applied to the regular data cubes, which limits the performance. To remedy this issue, in this paper we propose a novel irregular tensor low-rank representation model. We first segment the HSI data into several irregular homogeneous regions. Then, we propose a novel irregular tensor low-rank representation method that can efficiently model the irregular 3D cubes. We further use a non-convex nuclear norm to pursue the low-rankness and introduce a negative global low-rank term that improves global consistency. This proposed model is finally formulated as a convex-concave optimization problem and solved by alternative augmented Lagrangian method. Through experiments on four public datasets, the proposed method outperforms the existing low-rank based HSI methods significantly. Code is available at: https://github.com/hb-studying/ITLRR.

Via

Access Paper or Ask Questions

What If the Input is Expanded in OOD Detection?

Oct 24, 2024

Boxuan Zhang, Jianing Zhu, Zengmao Wang, Tongliang Liu, Bo Du, Bo Han

Figure 1 for What If the Input is Expanded in OOD Detection?

Figure 2 for What If the Input is Expanded in OOD Detection?

Figure 3 for What If the Input is Expanded in OOD Detection?

Figure 4 for What If the Input is Expanded in OOD Detection?

Abstract:Out-of-distribution (OOD) detection aims to identify OOD inputs from unknown classes, which is important for the reliable deployment of machine learning models in the open world. Various scoring functions are proposed to distinguish it from in-distribution (ID) data. However, existing methods generally focus on excavating the discriminative information from a single input, which implicitly limits its representation dimension. In this work, we introduce a novel perspective, i.e., employing different common corruptions on the input space, to expand that. We reveal an interesting phenomenon termed confidence mutation, where the confidence of OOD data can decrease significantly under the corruptions, while the ID data shows a higher confidence expectation considering the resistance of semantic features. Based on that, we formalize a new scoring method, namely, Confidence aVerage (CoVer), which can capture the dynamic differences by simply averaging the scores obtained from different corrupted inputs and the original ones, making the OOD and ID distributions more separable in detection tasks. Extensive experiments and analyses have been conducted to understand and verify the effectiveness of CoVer. The code is publicly available at: https://github.com/tmlr-group/CoVer.

* accepted by NeurIPS 2024

Via

Access Paper or Ask Questions

Mind the Gap Between Prototypes and Images in Cross-domain Finetuning

Oct 16, 2024

Hongduan Tian, Feng Liu, Zhanke Zhou, Tongliang Liu, Chengqi Zhang, Bo Han

Figure 1 for Mind the Gap Between Prototypes and Images in Cross-domain Finetuning

Figure 2 for Mind the Gap Between Prototypes and Images in Cross-domain Finetuning

Figure 3 for Mind the Gap Between Prototypes and Images in Cross-domain Finetuning

Figure 4 for Mind the Gap Between Prototypes and Images in Cross-domain Finetuning

Abstract:In cross-domain few-shot classification (CFC), recent works mainly focus on adapting a simple transformation head on top of a frozen pre-trained backbone with few labeled data to project embeddings into a task-specific metric space where classification can be performed by measuring similarities between image instance and prototype representations. Technically, an assumption implicitly adopted in such a framework is that the prototype and image instance embeddings share the same representation transformation. However, in this paper, we find that there naturally exists a gap, which resembles the modality gap, between the prototype and image instance embeddings extracted from the frozen pre-trained backbone, and simply applying the same transformation during the adaptation phase constrains exploring the optimal representations and shrinks the gap between prototype and image representations. To solve this problem, we propose a simple yet effective method, contrastive prototype-image adaptation (CoPA), to adapt different transformations respectively for prototypes and images similarly to CLIP by treating prototypes as text prompts. Extensive experiments on Meta-Dataset demonstrate that CoPA achieves the state-of-the-art performance more efficiently. Meanwhile, further analyses also indicate that CoPA can learn better representation clusters, enlarge the gap, and achieve minimal validation loss at the enlarged gap.

Via

Access Paper or Ask Questions

Knowledge-Enhanced Facial Expression Recognition with Emotional-to-Neutral Transformation

Sep 13, 2024

Hangyu Li, Yihan Xu, Jiangchao Yao, Nannan Wang, Xinbo Gao, Bo Han

Figure 1 for Knowledge-Enhanced Facial Expression Recognition with Emotional-to-Neutral Transformation

Figure 2 for Knowledge-Enhanced Facial Expression Recognition with Emotional-to-Neutral Transformation

Figure 3 for Knowledge-Enhanced Facial Expression Recognition with Emotional-to-Neutral Transformation

Figure 4 for Knowledge-Enhanced Facial Expression Recognition with Emotional-to-Neutral Transformation

Abstract:Existing facial expression recognition (FER) methods typically fine-tune a pre-trained visual encoder using discrete labels. However, this form of supervision limits to specify the emotional concept of different facial expressions. In this paper, we observe that the rich knowledge in text embeddings, generated by vision-language models, is a promising alternative for learning discriminative facial expression representations. Inspired by this, we propose a novel knowledge-enhanced FER method with an emotional-to-neutral transformation. Specifically, we formulate the FER problem as a process to match the similarity between a facial expression representation and text embeddings. Then, we transform the facial expression representation to a neutral representation by simulating the difference in text embeddings from textual facial expression to textual neutral. Finally, a self-contrast objective is introduced to pull the facial expression representation closer to the textual facial expression, while pushing it farther from the neutral representation. We conduct evaluation with diverse pre-trained visual encoders including ResNet-18 and Swin-T on four challenging facial expression datasets. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art FER methods. The code will be publicly available.

Via

Access Paper or Ask Questions

Text-based Talking Video Editing with Cascaded Conditional Diffusion

Jul 20, 2024

Bo Han, Heqing Zou, Haoyang Li, Guangcong Wang, Chng Eng Siong

Figure 1 for Text-based Talking Video Editing with Cascaded Conditional Diffusion

Figure 2 for Text-based Talking Video Editing with Cascaded Conditional Diffusion

Figure 3 for Text-based Talking Video Editing with Cascaded Conditional Diffusion

Figure 4 for Text-based Talking Video Editing with Cascaded Conditional Diffusion

Abstract:Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging because of \textbf{1)} generalizable talking-face representation, \textbf{2)} seamless audio-visual transitions, and \textbf{3)} identity-preserved talking faces. Previous works either require minutes of talking-face video training data and expensive test-time optimization for customized talking video editing or directly generate a video sequence without considering in-context information, leading to a poor generalizable representation, or incoherent transitions, or even inconsistent identity. In this paper, we propose an efficient cascaded conditional diffusion-based framework, which consists of two stages: audio to dense-landmark motion and motion to video. \textit{\textbf{In the first stage}}, we first propose a dynamic weighted in-context diffusion module to synthesize dense-landmark motions given an edited audio. \textit{\textbf{In the second stage}}, we introduce a warping-guided conditional diffusion module. The module first interpolates between the start and end frames of the editing interval to generate smooth intermediate frames. Then, with the help of the audio-to-dense motion images, these intermediate frames are warped to obtain coarse intermediate frames. Conditioned on the warped intermedia frames, a diffusion model is adopted to generate detailed and high-resolution target frames, which guarantees coherent and identity-preserved transitions. The cascaded conditional diffusion model decomposes the complex talking editing task into two flexible generation tasks, which provides a generalizable talking-face representation, seamless audio-visual transitions, and identity-preserved faces on a small dataset. Experiments show the effectiveness and superiority of the proposed method.

Via

Access Paper or Ask Questions

Robust Learning under Hybrid Noise

Jul 04, 2024

Yang Wei, Shuo Chen, Shanshan Ye, Bo Han, Chen Gong

Abstract:Feature noise and label noise are ubiquitous in practical scenarios, which pose great challenges for training a robust machine learning model. Most previous approaches usually deal with only a single problem of either feature noise or label noise. However, in real-world applications, hybrid noise, which contains both feature noise and label noise, is very common due to the unreliable data collection and annotation processes. Although some results have been achieved by a few representation learning based attempts, this issue is still far from being addressed with promising performance and guaranteed theoretical analyses. To address the challenge, we propose a novel unified learning framework called "Feature and Label Recovery" (FLR) to combat the hybrid noise from the perspective of data recovery, where we concurrently reconstruct both the feature matrix and the label matrix of input data. Specifically, the clean feature matrix is discovered by the low-rank approximation, and the ground-truth label matrix is embedded based on the recovered features with a nuclear norm regularization. Meanwhile, the feature noise and label noise are characterized by their respective adaptive matrix norms to satisfy the corresponding maximum likelihood. As this framework leads to a non-convex optimization problem, we develop the non-convex Alternating Direction Method of Multipliers (ADMM) with the convergence guarantee to solve our learning objective. We also provide the theoretical analysis to show that the generalization error of FLR can be upper-bounded in the presence of hybrid noise. Experimental results on several typical benchmark datasets clearly demonstrate the superiority of our proposed method over the state-of-the-art robust learning approaches for various noises.

Via

Access Paper or Ask Questions

Unlearning with Control: Assessing Real-world Utility for Large Language Model Unlearning

Jun 13, 2024

Qizhou Wang, Bo Han, Puning Yang, Jianing Zhu, Tongliang Liu, Masashi Sugiyama

Figure 1 for Unlearning with Control: Assessing Real-world Utility for Large Language Model Unlearning

Figure 2 for Unlearning with Control: Assessing Real-world Utility for Large Language Model Unlearning

Figure 3 for Unlearning with Control: Assessing Real-world Utility for Large Language Model Unlearning

Figure 4 for Unlearning with Control: Assessing Real-world Utility for Large Language Model Unlearning

Abstract:The compelling goal of eradicating undesirable data behaviors, while preserving usual model functioning, underscores the significance of machine unlearning within the domain of large language models (LLMs). Recent research has begun to approach LLM unlearning via gradient ascent (GA) -- increasing the prediction risk for those training strings targeted to be unlearned, thereby erasing their parameterized responses. Despite their simplicity and efficiency, we suggest that GA-based methods face the propensity towards excessive unlearning, resulting in various undesirable model behaviors, such as catastrophic forgetting, that diminish their practical utility. In this paper, we suggest a set of metrics that can capture multiple facets of real-world utility and propose several controlling methods that can regulate the extent of excessive unlearning. Accordingly, we suggest a general framework to better reflect the practical efficacy of various unlearning methods -- we begin by controlling the unlearning procedures/unlearned models such that no excessive unlearning occurs and follow by the evaluation for unlearning efficacy. Our experimental analysis on established benchmarks revealed that GA-based methods are far from perfect in practice, as strong unlearning is at the high cost of hindering the model utility. We conclude that there is still a long way towards practical and effective LLM unlearning, and more efforts are required in this field.

Via

Access Paper or Ask Questions