Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Yang Yang

Department of Biostatistics and Bioinformatics, Duke University, Durham, USA

The Solution for Temporal Action Localisation Task of Perception Test Challenge 2024

Oct 08, 2024

Yinan Han, Qingyuan Jiang, Hongming Mei, Yang Yang, Jinhui Tang

Figure 1 for The Solution for Temporal Action Localisation Task of Perception Test Challenge 2024

Figure 2 for The Solution for Temporal Action Localisation Task of Perception Test Challenge 2024

Figure 3 for The Solution for Temporal Action Localisation Task of Perception Test Challenge 2024

Figure 4 for The Solution for Temporal Action Localisation Task of Perception Test Challenge 2024

Abstract:This report presents our method for Temporal Action Localisation (TAL), which focuses on identifying and classifying actions within specific time intervals throughout a video sequence. We employ a data augmentation technique by expanding the training dataset using overlapping labels from the Something-SomethingV2 dataset, enhancing the model's ability to generalize across various action classes. For feature extraction, we utilize state-of-the-art models, including UMT, VideoMAEv2 for video features, and BEATs and CAV-MAE for audio features. Our approach involves training both multimodal (video and audio) and unimodal (video only) models, followed by combining their predictions using the Weighted Box Fusion (WBF) method. This fusion strategy ensures robust action localisation. our overall approach achieves a score of 0.5498, securing first place in the competition.

Via

Access Paper or Ask Questions

BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data

Oct 01, 2024

Xuwu Wang, Qiwen Cui, Yunzhe Tao, Yiran Wang, Ziwei Chai, Xiaotian Han, Boyi Liu, Jianbo Yuan, Jing Su, Guoyin Wang(+9 more)

Figure 1 for BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data

Figure 2 for BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data

Figure 3 for BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data

Figure 4 for BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data

Abstract:Large language models (LLMs) have become increasingly pivotal across various domains, especially in handling complex data types. This includes structured data processing, as exemplified by ChartQA and ChatGPT-Ada, and multimodal unstructured data processing as seen in Visual Question Answering (VQA). These areas have attracted significant attention from both industry and academia. Despite this, there remains a lack of unified evaluation methodologies for these diverse data handling scenarios. In response, we introduce BabelBench, an innovative benchmark framework that evaluates the proficiency of LLMs in managing multimodal multistructured data with code execution. BabelBench incorporates a dataset comprising 247 meticulously curated problems that challenge the models with tasks in perception, commonsense reasoning, logical reasoning, and so on. Besides the basic capabilities of multimodal understanding, structured data processing as well as code generation, these tasks demand advanced capabilities in exploration, planning, reasoning and debugging. Our experimental findings on BabelBench indicate that even cutting-edge models like ChatGPT 4 exhibit substantial room for improvement. The insights derived from our comprehensive analysis offer valuable guidance for future research within the community. The benchmark data can be found at https://github.com/FFD8FFE/babelbench.

Via

Access Paper or Ask Questions

Solution for OOD-CV Workshop SSB Challenge 2024 (Open-Set Recognition Track)

Sep 30, 2024

Mingxu Feng, Dian Chao, Peng Zheng, Yang Yang

Figure 1 for Solution for OOD-CV Workshop SSB Challenge 2024 (Open-Set Recognition Track)

Figure 2 for Solution for OOD-CV Workshop SSB Challenge 2024 (Open-Set Recognition Track)

Figure 3 for Solution for OOD-CV Workshop SSB Challenge 2024 (Open-Set Recognition Track)

Figure 4 for Solution for OOD-CV Workshop SSB Challenge 2024 (Open-Set Recognition Track)

Abstract:This report provides a detailed description of the method we explored and proposed in the OSR Challenge at the OOD-CV Workshop during ECCV 2024. The challenge required identifying whether a test sample belonged to the semantic classes of a classifier's training set, a task known as open-set recognition (OSR). Using the Semantic Shift Benchmark (SSB) for evaluation, we focused on ImageNet1k as the in-distribution (ID) dataset and a subset of ImageNet21k as the out-of-distribution (OOD) dataset.To address this, we proposed a hybrid approach, experimenting with the fusion of various post-hoc OOD detection techniques and different Test-Time Augmentation (TTA) strategies. Additionally, we evaluated the impact of several base models on the final performance. Our best-performing method combined Test-Time Augmentation with the post-hoc OOD techniques, achieving a strong balance between AUROC and FPR95 scores. Our approach resulted in AUROC: 79.77 (ranked 5th) and FPR95: 61.44 (ranked 2nd), securing second place in the overall competition.

Via

Access Paper or Ask Questions

Solution for Temporal Sound Localisation Task of ECCV Second Perception Test Challenge 2024

Sep 29, 2024

Haowei Gu, Weihao Zhu, Yang Yang

Figure 1 for Solution for Temporal Sound Localisation Task of ECCV Second Perception Test Challenge 2024

Figure 2 for Solution for Temporal Sound Localisation Task of ECCV Second Perception Test Challenge 2024

Abstract:This report proposes an improved method for the Temporal Sound Localisation (TSL) task, which localizes and classifies the sound events occurring in the video according to a predefined set of sound classes. The champion solution from last year's first competition has explored the TSL by fusing audio and video modalities with the same weight. Considering the TSL task aims to localize sound events, we conduct relevant experiments that demonstrated the superiority of sound features (Section 3). Based on our findings, to enhance audio modality features, we employ various models to extract audio features, such as InterVideo, CaVMAE, and VideoMAE models. Our approach ranks first in the final test with a score of 0.4925.

Via

Access Paper or Ask Questions

A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Sep 28, 2024

Houjian Yu, Mingen Li, Alireza Rezazadeh, Yang Yang, Changhyun Choi

Figure 1 for A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Figure 2 for A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Figure 3 for A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Figure 4 for A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Abstract:The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large Language Models (MLLMs) have shown promising results, their extensive computation and data demands limit the feasibility of local deployment and customization. To address this, we propose a novel CLIP-based multimodal parameter-efficient tuning (PET) framework designed for three language-guided object grounding and grasping tasks: (1) Referring Expression Segmentation (RES), (2) Referring Grasp Synthesis (RGS), and (3) Referring Grasp Affordance (RGA). Our approach introduces two key innovations: a bi-directional vision-language adapter that aligns multimodal inputs for pixel-level language understanding and a depth fusion branch that incorporates geometric cues to facilitate robot grasping predictions. Experiment results demonstrate superior performance in the RES object grounding task compared with existing CLIP-based full-model tuning or PET approaches. In the RGS and RGA tasks, our model not only effectively interprets object attributes based on simple language descriptions but also shows strong potential for comprehending complex spatial reasoning scenarios, such as multiple identical objects present in the workspace.

* This work has been submitted to ICRA 2025

Via

Access Paper or Ask Questions

SVFit: Parameter-Efficient Fine-Tuning of Large Pre-Trained Models Using Singular Values

Sep 09, 2024

Chengwei Sun, Jiwei Wei, Yujia Wu, Yiming Shi, Shiyuan He, Zeyu Ma, Ning Xie, Yang Yang

Figure 1 for SVFit: Parameter-Efficient Fine-Tuning of Large Pre-Trained Models Using Singular Values

Figure 2 for SVFit: Parameter-Efficient Fine-Tuning of Large Pre-Trained Models Using Singular Values

Figure 3 for SVFit: Parameter-Efficient Fine-Tuning of Large Pre-Trained Models Using Singular Values

Figure 4 for SVFit: Parameter-Efficient Fine-Tuning of Large Pre-Trained Models Using Singular Values

Abstract:Large pre-trained models (LPMs) have demonstrated exceptional performance in diverse natural language processing and computer vision tasks. However, fully fine-tuning these models poses substantial memory challenges, particularly in resource-constrained environments. Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, mitigate this issue by adjusting only a small subset of parameters. Nevertheless, these methods typically employ random initialization for low-rank matrices, which can lead to inefficiencies in gradient descent and diminished generalizability due to suboptimal starting points. To address these limitations, we propose SVFit, a novel PEFT approach that leverages singular value decomposition (SVD) to initialize low-rank matrices using critical singular values as trainable parameters. Specifically, SVFit performs SVD on the pre-trained weight matrix to obtain the best rank-r approximation matrix, emphasizing the most critical singular values that capture over 99% of the matrix's information. These top-r singular values are then used as trainable parameters to scale the fundamental subspaces of the matrix, facilitating rapid domain adaptation. Extensive experiments across various pre-trained models in natural language understanding, text-to-image generation, and image classification tasks reveal that SVFit outperforms LoRA while requiring 16 times fewer trainable parameters.

Via

Access Paper or Ask Questions

Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance

Sep 06, 2024

RenMing Huang, Shaochong Liu, Yunqiang Pei, Peng Wang, Guoqing Wang, Yang Yang, Hengtao Shen

Figure 1 for Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance

Figure 2 for Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance

Figure 3 for Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance

Figure 4 for Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance

Abstract:In this work, we address the challenging problem of long-horizon goal-reaching policy learning from non-expert, action-free observation data. Unlike fully labeled expert data, our data is more accessible and avoids the costly process of action labeling. Additionally, compared to online learning, which often involves aimless exploration, our data provides useful guidance for more efficient exploration. To achieve our goal, we propose a novel subgoal guidance learning strategy. The motivation behind this strategy is that long-horizon goals offer limited guidance for efficient exploration and accurate state transition. We develop a diffusion strategy-based high-level policy to generate reasonable subgoals as waypoints, preferring states that more easily lead to the final goal. Additionally, we learn state-goal value functions to encourage efficient subgoal reaching. These two components naturally integrate into the off-policy actor-critic framework, enabling efficient goal attainment through informative exploration. We evaluate our method on complex robotic navigation and manipulation tasks, demonstrating a significant performance advantage over existing methods. Our ablation study further shows that our method is robust to observation data with various corruptions.

* Accepted to CoRL 2024

Via

Access Paper or Ask Questions

Reinforcement-Learning-Enabled Beam Alignment for Water-Air Direct Optical Wireless Communications

Sep 05, 2024

Jiayue Liu, Tianqi Mao, Dongxuan He, Yang Yang, Zhen Gao, Dezhi Zheng, Jun Zhang

Figure 1 for Reinforcement-Learning-Enabled Beam Alignment for Water-Air Direct Optical Wireless Communications

Figure 2 for Reinforcement-Learning-Enabled Beam Alignment for Water-Air Direct Optical Wireless Communications

Figure 3 for Reinforcement-Learning-Enabled Beam Alignment for Water-Air Direct Optical Wireless Communications

Figure 4 for Reinforcement-Learning-Enabled Beam Alignment for Water-Air Direct Optical Wireless Communications

Abstract:The escalating interests on underwater exploration/reconnaissance applications have motivated high-rate data transmission from underwater to airborne relaying platforms, especially under high-sea scenarios. Thanks to its broad bandwidth and superior confidentiality, Optical wireless communication has become one promising candidate for water-air transmission. However, the optical signals inevitably suffer from deviations when crossing the highly-dynamic water-air interfaces in the absence of relaying ships/buoys. To address the issue, this article proposes one novel beam alignment strategy based on deep reinforcement learning (DRL) for water-air direct optical wireless communications. Specifically, the dynamic water-air interface is mathematically modeled using sea-wave spectrum analysis, followed by characterization of the propagation channel with ray-tracing techniques. Then the deep deterministic policy gradient (DDPG) scheme is introduced for DRL-based transceiving beam alignment. A logarithm-exponential (LE) nonlinear reward function with respect to the received signal strength is designed for high-resolution rewarding between different actions. Simulation results validate the superiority of the proposed DRL-based beam alignment scheme.

* 6 pages, 6 figures, published to ICCC 2024

Via

Access Paper or Ask Questions

RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

Sep 05, 2024

Benzhi Wang, Jingkai Zhou, Jingqi Bai, Yang Yang, Weihua Chen, Fan Wang, Zhen Lei

Figure 1 for RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

Figure 2 for RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

Figure 3 for RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

Figure 4 for RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

Abstract:In recent years, diffusion models have revolutionized visual generation, outperforming traditional frameworks like Generative Adversarial Networks (GANs). However, generating images of humans with realistic semantic parts, such as hands and faces, remains a significant challenge due to their intricate structural complexity. To address this issue, we propose a novel post-processing solution named RealisHuman. The RealisHuman framework operates in two stages. First, it generates realistic human parts, such as hands or faces, using the original malformed parts as references, ensuring consistent details with the original image. Second, it seamlessly integrates the rectified human parts back into their corresponding positions by repainting the surrounding areas to ensure smooth and realistic blending. The RealisHuman framework significantly enhances the realism of human generation, as demonstrated by notable improvements in both qualitative and quantitative metrics. Code is available at https://github.com/Wangbenzhi/RealisHuman.

Via

Access Paper or Ask Questions

Brant-X: A Unified Physiological Signal Alignment Framework

Aug 28, 2024

Daoze Zhang, Zhizhang Yuan, Junru Chen, Kerui Chen, Yang Yang

Figure 1 for Brant-X: A Unified Physiological Signal Alignment Framework

Figure 2 for Brant-X: A Unified Physiological Signal Alignment Framework

Figure 3 for Brant-X: A Unified Physiological Signal Alignment Framework

Figure 4 for Brant-X: A Unified Physiological Signal Alignment Framework

Abstract:Physiological signals serve as indispensable clues for understanding various physiological states of human bodies. Most existing works have focused on a single type of physiological signals for a range of application scenarios. However, as the body is a holistic biological system, the inherent interconnection among various physiological data should not be neglected. In particular, given the brain's role as the control center for vital activities, electroencephalogram (EEG) exhibits significant correlations with other physiological signals. Therefore, the correlation between EEG and other physiological signals holds potential to improve performance in various scenarios. Nevertheless, achieving this goal is still constrained by several challenges: the scarcity of simultaneously collected physiological data, the differences in correlations between various signals, and the correlation differences between various tasks. To address these issues, we propose a unified physiological signal alignment framework, Brant-X, to model the correlation between EEG and other signals. Our approach (1) employs the EEG foundation model to data-efficiently transfer the rich knowledge in EEG to other physiological signals, and (2) introduces the two-level alignment to fully align the semantics of EEG and other signals from different semantic scales. In the experiments, Brant-X achieves state-of-the-art performance compared with task-agnostic and task-specific baselines on various downstream tasks in diverse scenarios, including sleep stage classification, emotion recognition, freezing of gaits detection, and eye movement communication. Moreover, the analysis on the arrhythmia detection task and the visualization in case study further illustrate the effectiveness of Brant-X in the knowledge transfer from EEG to other physiological signals. The model's homepage is at https://github.com/zjunet/Brant-X/.

* SIGKDD 2024
* Accepted by SIGKDD 2024

Via

Access Paper or Ask Questions