Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Andrew Rouditchenko

Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

Dec 08, 2021

Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogerio Feris, David Harwath, James Glass, Hilde Kuehne

Figure 1 for Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

Figure 2 for Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

Figure 3 for Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

Figure 4 for Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

Abstract:Multi-modal learning from video data has seen increased attention recently as it allows to train semantically meaningful embeddings without human annotation enabling tasks like zero-shot retrieval and classification. In this work, we present a multi-modal, modality agnostic fusion transformer approach that learns to exchange information between multiple modalities, such as video, audio, and text, and integrate them into a joined multi-modal representation to obtain an embedding that aggregates multi-modal temporal information. We propose to train the system with a combinatorial loss on everything at once, single modalities as well as pairs of modalities, explicitly leaving out any add-ons such as position or modality encoding. At test time, the resulting model can process and fuse any number of input modalities. Moreover, the implicit properties of the transformer allow to process inputs of different lengths. To evaluate the proposed approach, we train the model on the large scale HowTo100M dataset and evaluate the resulting embedding space on four challenging benchmark datasets obtaining state-of-the-art results in zero-shot video retrieval and zero-shot video action localization.

Via

Access Paper or Ask Questions

Routing with Self-Attention for Multimodal Capsule Networks

Dec 01, 2021

Kevin Duarte, Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Samuel Thomas, Alexander Liu, David Harwath, James Glass, Hilde Kuehne, Mubarak Shah

Figure 1 for Routing with Self-Attention for Multimodal Capsule Networks

Figure 2 for Routing with Self-Attention for Multimodal Capsule Networks

Figure 3 for Routing with Self-Attention for Multimodal Capsule Networks

Figure 4 for Routing with Self-Attention for Multimodal Capsule Networks

Abstract:The task of multimodal learning has seen a growing interest recently as it allows for training neural architectures based on different modalities such as vision, text, and audio. One challenge in training such models is that they need to jointly learn semantic concepts and their relationships across different input representations. Capsule networks have been shown to perform well in context of capturing the relation between low-level input features and higher-level concepts. However, capsules have so far mainly been used only in small-scale fully supervised settings due to the resource demand of conventional routing algorithms. We present a new multimodal capsule network that allows us to leverage the strength of capsules in the context of a multimodal learning framework on large amounts of video data. To adapt the capsules to large-scale input data, we propose a novel routing by self-attention mechanism that selects relevant capsules which are then used to generate a final joint multimodal feature representation. This allows not only for robust training with noisy video data, but also to scale up the size of the capsule network compared to traditional routing methods while still being computationally efficient. We evaluate the proposed architecture by pretraining it on a large-scale multimodal video dataset and applying it on four datasets in two challenging downstream tasks. Results show that the proposed multimodal capsule network is not only able to improve results compared to other routing techniques, but also achieves competitive performance on the task of multimodal learning.

Via

Access Paper or Ask Questions

Cascaded Multilingual Audio-Visual Learning from Videos

Nov 08, 2021

Andrew Rouditchenko, Angie Boggust, David Harwath, Samuel Thomas, Hilde Kuehne, Brian Chen, Rameswar Panda, Rogerio Feris, Brian Kingsbury, Michael Picheny(+1 more)

Figure 1 for Cascaded Multilingual Audio-Visual Learning from Videos

Figure 2 for Cascaded Multilingual Audio-Visual Learning from Videos

Figure 3 for Cascaded Multilingual Audio-Visual Learning from Videos

Figure 4 for Cascaded Multilingual Audio-Visual Learning from Videos

Abstract:In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to visual content after training on a large-scale dataset of videos, but they were only trained and evaluated on videos in English. To learn multilingual audio-visual representations, we propose a cascaded approach that leverages a model trained on English videos and applies it to audio-visual data in other languages, such as Japanese videos. With our cascaded approach, we show an improvement in retrieval performance of nearly 10x compared to training on the Japanese videos solely. We also apply the model trained on English videos to Japanese and Hindi spoken captions of images, achieving state-of-the-art performance.

* Presented at Interspeech 2021. This version contains updated results using the YouCook-Japanese dataset

Via

Access Paper or Ask Questions

Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset

Oct 14, 2021

Ian Palmer, Andrew Rouditchenko, Andrei Barbu, Boris Katz, James Glass

Figure 1 for Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset

Figure 2 for Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset

Figure 3 for Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset

Figure 4 for Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset

Abstract:Visually-grounded spoken language datasets can enable models to learn cross-modal correspondences with very weak supervision. However, modern audio-visual datasets contain biases that undermine the real-world performance of models trained on that data. We introduce Spoken ObjectNet, which is designed to remove some of these biases and provide a way to better evaluate how effectively models will perform in real-world scenarios. This dataset expands upon ObjectNet, which is a bias-controlled image dataset that features similar image classes to those present in ImageNet. We detail our data collection pipeline, which features several methods to improve caption quality, including automated language model checks. Lastly, we show baseline results on image retrieval and audio retrieval tasks. These results show that models trained on other datasets and then evaluated on Spoken ObjectNet tend to perform poorly due to biases in other datasets that the models have learned. We also show evidence that the performance decrease is due to the dataset controls, and not the transfer setting.

* Presented at Interspeech 2021. This version contains additional experiments on the Spoken ObjectNet test set

Via

Access Paper or Ask Questions

Cross-Modal Discrete Representation Learning

Jun 10, 2021

Alexander H. Liu, SouYoung Jin, Cheng-I Jeff Lai, Andrew Rouditchenko, Aude Oliva, James Glass

Figure 1 for Cross-Modal Discrete Representation Learning

Figure 2 for Cross-Modal Discrete Representation Learning

Figure 3 for Cross-Modal Discrete Representation Learning

Figure 4 for Cross-Modal Discrete Representation Learning

Abstract:Recent advances in representation learning have demonstrated an ability to represent information from different modalities such as video, text, and audio in a single high-level embedding vector. In this work we present a self-supervised learning framework that is able to learn a representation that captures finer levels of granularity across different modalities such as concepts or events represented by visual objects or spoken words. Our framework relies on a discretized embedding space created via vector quantization that is shared across different modalities. Beyond the shared embedding space, we propose a Cross-Modal Code Matching objective that forces the representations from different views (modalities) to have a similar distribution over the discrete embedding space such that cross-modal objects/actions localization can be performed without direct supervision. In our experiments we show that the proposed discretized multi-modal fine-grained representation (e.g., pixel/word/frame) can complement high-level summary representations (e.g., video/sentence/waveform) for improved performance on cross-modal retrieval tasks. We also observe that the discretized representation uses individual clusters to represent the same semantic concept across modalities.

* Preprint

Via

Access Paper or Ask Questions

Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

May 05, 2021

Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie Boggust, Rameswar Panda, Brian Kingsbury, Rogerio Feris, David Harwath(+3 more)

Figure 1 for Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Figure 2 for Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Figure 3 for Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Figure 4 for Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Abstract:Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a self-supervised training framework that learns a common multimodal embedding space that, in addition to sharing representations across different modalities, enforces a grouping of semantically similar instances. To this end, we extend the concept of instance-level contrastive learning with a multimodal clustering step in the training pipeline to capture semantic similarities across modalities. The resulting embedding space enables retrieval of samples across all modalities, even from unseen datasets and different domains. To evaluate our approach, we train our model on the HowTo100M dataset and evaluate its zero-shot retrieval capabilities in two challenging domains, namely text-to-video retrieval, and temporal action localization, showing state-of-the-art results on four different datasets.

Via

Access Paper or Ask Questions

AVLnet: Learning Audio-Visual Language Representations from Instructional Videos

Jun 16, 2020

Andrew Rouditchenko, Angie Boggust, David Harwath, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba(+1 more)

Figure 1 for AVLnet: Learning Audio-Visual Language Representations from Instructional Videos

Figure 2 for AVLnet: Learning Audio-Visual Language Representations from Instructional Videos

Figure 3 for AVLnet: Learning Audio-Visual Language Representations from Instructional Videos

Figure 4 for AVLnet: Learning Audio-Visual Language Representations from Instructional Videos

Abstract:Current methods for learning visually grounded language from videos often rely on time-consuming and expensive data collection, such as human annotated textual summaries or machine generated automatic speech recognition transcripts. In this work, we introduce Audio-Video Language Network (AVLnet), a self-supervised network that learns a shared audio-visual embedding space directly from raw video inputs. We circumvent the need for annotation and instead learn audio-visual language representations directly from randomly segmented video clips and their raw audio waveforms. We train AVLnet on publicly available instructional videos and evaluate our model on video clip and language retrieval tasks on three video datasets. Our proposed model outperforms several state-of-the-art text-video baselines by up to 11.8% in a video clip retrieval task, despite operating on the raw audio instead of manually annotated text captions. Further, we show AVLnet is capable of integrating textual information, increasing its modularity and improving performance by up to 20.3% on the video clip retrieval task. Finally, we perform analysis of AVLnet's learned representations, showing our model has learned to relate visual objects with salient words and natural sounds.

Via

Access Paper or Ask Questions

Label-efficient audio classification through multitask learning and self-supervision

Oct 19, 2019

Tyler Lee, Ting Gong, Suchismita Padhy, Andrew Rouditchenko, Anthony Ndirango

Figure 1 for Label-efficient audio classification through multitask learning and self-supervision

Figure 2 for Label-efficient audio classification through multitask learning and self-supervision

Figure 3 for Label-efficient audio classification through multitask learning and self-supervision

Figure 4 for Label-efficient audio classification through multitask learning and self-supervision

Abstract:While deep learning has been incredibly successful in modeling tasks with large, carefully curated labeled datasets, its application to problems with limited labeled data remains a challenge. The aim of the present work is to improve the label efficiency of large neural networks operating on audio data through a combination of multitask learning and self-supervised learning on unlabeled data. We trained an end-to-end audio feature extractor based on WaveNet that feeds into simple, yet versatile task-specific neural networks. We describe several easily implemented self-supervised learning tasks that can operate on any large, unlabeled audio corpus. We demonstrate that, in scenarios with limited labeled training data, one can significantly improve the performance of three different supervised classification tasks individually by up to 6% through simultaneous training with these additional self-supervised tasks. We also show that incorporating data augmentation into our multitask setting leads to even further gains in performance.

* Presented at ICLR 2019 Limited Labeled Data (LLD) Workshop

Via

Access Paper or Ask Questions

Self-Supervised Audio-Visual Co-Segmentation

Apr 18, 2019

Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, Antonio Torralba

Figure 1 for Self-Supervised Audio-Visual Co-Segmentation

Figure 2 for Self-Supervised Audio-Visual Co-Segmentation

Figure 3 for Self-Supervised Audio-Visual Co-Segmentation

Figure 4 for Self-Supervised Audio-Visual Co-Segmentation

Abstract:Segmenting objects in images and separating sound sources in audio are challenging tasks, in part because traditional approaches require large amounts of labeled data. In this paper we develop a neural network model for visual object segmentation and sound source separation that learns from natural videos through self-supervision. The model is an extension of recently proposed work that maps image pixels to sounds. Here, we introduce a learning approach to disentangle concepts in the neural networks, and assign semantic categories to network feature channels to enable independent image segmentation and sound source separation after audio-visual training on videos. Our evaluations show that the disentangled model outperforms several baselines in semantic segmentation and sound source separation.

* Accepted to ICASSP 2019

Via

Access Paper or Ask Questions

The Sound of Pixels

Oct 14, 2018

Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, Antonio Torralba

Abstract:We introduce PixelPlayer, a system that, by leveraging large amounts of unlabeled videos, learns to locate image regions which produce sounds and separate the input sounds into a set of components that represents the sound from each pixel. Our approach capitalizes on the natural synchronization of the visual and audio modalities to learn models that jointly parse sounds and images, without requiring additional manual supervision. Experimental results on a newly collected MUSIC dataset show that our proposed Mix-and-Separate framework outperforms several baselines on source separation. Qualitative results suggest our model learns to ground sounds in vision, enabling applications such as independently adjusting the volume of sound sources.

Via

Access Paper or Ask Questions