Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Sriram Ganapathy

STAB: Speech Tokenizer Assessment Benchmark

Sep 04, 2024

Shikhar Vashishth, Harman Singh, Shikhar Bharadwaj, Sriram Ganapathy, Chulayuth Asawaroengchai, Kartik Audhkhasi, Andrew Rosenberg, Ankur Bapna, Bhuvana Ramabhadran

Figure 1 for STAB: Speech Tokenizer Assessment Benchmark

Figure 2 for STAB: Speech Tokenizer Assessment Benchmark

Figure 3 for STAB: Speech Tokenizer Assessment Benchmark

Figure 4 for STAB: Speech Tokenizer Assessment Benchmark

Abstract:Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs). Currently, while several speech tokenizers have been proposed, there is ambiguity regarding the properties that are desired from a tokenizer for specific downstream tasks and its overall generalizability. Evaluating the performance of tokenizers across different downstream tasks is a computationally intensive effort that poses challenges for scalability. To circumvent this requirement, we present STAB (Speech Tokenizer Assessment Benchmark), a systematic evaluation framework designed to assess speech tokenizers comprehensively and shed light on their inherent characteristics. This framework provides a deeper understanding of the underlying mechanisms of speech tokenization, thereby offering a valuable resource for expediting the advancement of future tokenizer models and enabling comparative analysis using a standardized benchmark. We evaluate the STAB metrics and correlate this with downstream task performance across a range of speech tasks and tokenizer choices.

* 5 pages

Via

Access Paper or Ask Questions

Improving Self-supervised Pre-training using Accent-Specific Codebooks

Jul 04, 2024

Darshan Prabhu, Abhishek Gupta, Omkar Nitsure, Preethi Jyothi, Sriram Ganapathy

Figure 1 for Improving Self-supervised Pre-training using Accent-Specific Codebooks

Figure 2 for Improving Self-supervised Pre-training using Accent-Specific Codebooks

Figure 3 for Improving Self-supervised Pre-training using Accent-Specific Codebooks

Figure 4 for Improving Self-supervised Pre-training using Accent-Specific Codebooks

Abstract:Speech accents present a serious challenge to the performance of state-of-the-art end-to-end Automatic Speech Recognition (ASR) systems. Even with self-supervised learning and pre-training of ASR models, accent invariance is seldom achieved. In this work, we propose an accent-aware adaptation technique for self-supervised learning that introduces a trainable set of accent-specific codebooks to the self-supervised architecture. These learnable codebooks enable the model to capture accent specific information during pre-training, that is further refined during ASR finetuning. On the Mozilla Common Voice dataset, our proposed approach outperforms all other accent-adaptation approaches on both seen and unseen English accents, with up to 9% relative reduction in word error rate (WER).

* Accepted to INTERSPEECH 2024

Via

Access Paper or Ask Questions

Towards the Next Frontier in Speech Representation Learning Using Disentanglement

Jul 02, 2024

Varun Krishna, Sriram Ganapathy

Figure 1 for Towards the Next Frontier in Speech Representation Learning Using Disentanglement

Figure 2 for Towards the Next Frontier in Speech Representation Learning Using Disentanglement

Figure 3 for Towards the Next Frontier in Speech Representation Learning Using Disentanglement

Figure 4 for Towards the Next Frontier in Speech Representation Learning Using Disentanglement

Abstract:The popular frameworks for self-supervised learning of speech representations have largely focused on frame-level masked prediction of speech regions. While this has shown promising downstream task performance for speech recognition and related tasks, this has largely ignored factors of speech that are encoded at coarser level, like characteristics of the speaker or channel that remain consistent through-out a speech utterance. In this work, we propose a framework for Learning Disentangled Self Supervised (termed as Learn2Diss) representations of speech, which consists of frame-level and an utterance-level encoder modules. The two encoders are initially learned independently, where the frame-level model is largely inspired by existing self supervision techniques, thereby learning pseudo-phonemic representations, while the utterance-level encoder is inspired by constrastive learning of pooled embeddings, thereby learning pseudo-speaker representations. The joint learning of these two modules consists of disentangling the two encoders using a mutual information based criterion. With several downstream evaluation experiments, we show that the proposed Learn2Diss achieves state-of-the-art results on a variety of tasks, with the frame-level encoder representations improving semantic tasks, while the utterance-level representations improve non-semantic tasks.

Via

Access Paper or Ask Questions

The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments

Jun 13, 2024

Shareef Babu Kalluri, Prachi Singh, Pratik Roy Chowdhuri, Apoorva Kulkarni, Shikha Baghel, Pradyoth Hegde, Swapnil Sontakke, Deepak K T, S. R. Mahadeva Prasanna, Deepu Vijayasenan(+1 more)

Figure 1 for The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments

Figure 2 for The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments

Figure 3 for The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments

Figure 4 for The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments

Abstract:The DIarization of SPeaker and LAnguage in Conversational Environments (DISPLACE) 2024 challenge is the second in the series of DISPLACE challenges, which involves tasks of speaker diarization (SD) and language diarization (LD) on a challenging multilingual conversational speech dataset. In the DISPLACE 2024 challenge, we also introduced the task of automatic speech recognition (ASR) on this dataset. The dataset containing 158 hours of speech, consisting of both supervised and unsupervised mono-channel far-field recordings, was released for LD and SD tracks. Further, 12 hours of close-field mono-channel recordings were provided for the ASR track conducted on 5 Indian languages. The details of the dataset, baseline systems and the leader board results are highlighted in this paper. We have also compared our baseline models and the team's performances on evaluation data of DISPLACE-2023 to emphasize the advancements made in this second version of the challenge.

* 5 pages, 3 figures, Interspeech 2024

Via

Access Paper or Ask Questions

Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization

Jan 23, 2024

Prachi Singh, Sriram Ganapathy

Figure 1 for Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization

Figure 2 for Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization

Figure 3 for Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization

Figure 4 for Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization

Abstract:Speaker diarization, the task of segmenting an audio recording based on speaker identity, constitutes an important speech pre-processing step for several downstream applications. The conventional approach to diarization involves multiple steps of embedding extraction and clustering, which are often optimized in an isolated fashion. While end-to-end diarization systems attempt to learn a single model for the task, they are often cumbersome to train and require large supervised datasets. In this paper, we propose an end-to-end supervised hierarchical clustering algorithm based on graph neural networks (GNN), called End-to-end Supervised HierARchical Clustering (E-SHARC). The E-SHARC approach uses front-end mel-filterbank features as input and jointly learns an embedding extractor and the GNN clustering module, performing representation learning, metric learning, and clustering with end-to-end optimization. Further, with additional inputs from an external overlap detector, the E-SHARC approach is capable of predicting the speakers in the overlapping speech regions. The experimental evaluation on several benchmark datasets like AMI, VoxConverse and DISPLACE, illustrates that the proposed E-SHARC framework improves significantly over the state-of-art diarization systems.

* 10 pages

Via

Access Paper or Ask Questions

Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement

Jan 09, 2024

Soumya Dutta, Sriram Ganapathy

Figure 1 for Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement

Figure 2 for Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement

Figure 3 for Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement

Figure 4 for Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement

Abstract:The problem of audio-to-audio (A2A) style transfer involves replacing the style features of the source audio with those from the target audio while preserving the content related attributes of the source audio. In this paper, we propose an efficient approach, termed as Zero-shot Emotion Style Transfer (ZEST), that allows the transfer of emotional content present in the given source audio with the one embedded in the target audio while retaining the speaker and speech content from the source. The proposed system builds upon decomposing speech into semantic tokens, speaker representations and emotion embeddings. Using these factors, we propose a framework to reconstruct the pitch contour of the given speech signal and train a decoder that reconstructs the speech signal. The model is trained using a self-supervision based reconstruction loss. During conversion, the emotion embedding is alone derived from the target audio, while rest of the factors are derived from the source audio. In our experiments, we show that, even without using parallel training data or labels from the source or target audio, we illustrate zero shot emotion transfer capabilities of the proposed ZEST model using objective and subjective quality evaluations.

* 5 pages, 3 figures, accepted at ICASSP 2024

Via

Access Paper or Ask Questions

LLM Augmented LLMs: Expanding Capabilities through Composition

Jan 04, 2024

Rachit Bansal, Bidisha Samanta, Siddharth Dalmia, Nitish Gupta, Shikhar Vashishth, Sriram Ganapathy, Abhishek Bapna, Prateek Jain, Partha Talukdar

Figure 1 for LLM Augmented LLMs: Expanding Capabilities through Composition

Figure 2 for LLM Augmented LLMs: Expanding Capabilities through Composition

Figure 3 for LLM Augmented LLMs: Expanding Capabilities through Composition

Figure 4 for LLM Augmented LLMs: Expanding Capabilities through Composition

Abstract:Foundational models with billions of parameters which have been trained on large corpora of data have demonstrated non-trivial skills in a variety of domains. However, due to their monolithic structure, it is challenging and expensive to augment them or impart new skills. On the other hand, due to their adaptation abilities, several new instances of these models are being trained towards new domains and tasks. In this work, we study the problem of efficient and practical composition of existing foundation models with more specific models to enable newer capabilities. To this end, we propose CALM -- Composition to Augment Language Models -- which introduces cross-attention between models to compose their representations and enable new capabilities. Salient features of CALM are: (i) Scales up LLMs on new tasks by 're-using' existing LLMs along with a few additional parameters and data, (ii) Existing model weights are kept intact, and hence preserves existing capabilities, and (iii) Applies to diverse domains and settings. We illustrate that augmenting PaLM2-S with a smaller model trained on low-resource languages results in an absolute improvement of up to 13\% on tasks like translation into English and arithmetic reasoning for low-resource languages. Similarly, when PaLM2-S is augmented with a code-specific model, we see a relative improvement of 40\% over the base model for code generation and explanation tasks -- on-par with fully fine-tuned counterparts.

* 17 pages, 2 figures, 8 tables

Via

Access Paper or Ask Questions

Summary of the DISPLACE Challenge 2023 -- DIarization of SPeaker and LAnguage in Conversational Environments

Nov 23, 2023

Shikha Baghel, Shreyas Ramoji, Somil Jain, Pratik Roy Chowdhuri, Prachi Singh, Deepu Vijayasenan, Sriram Ganapathy

Abstract:In multi-lingual societies, where multiple languages are spoken in a small geographic vicinity, informal conversations often involve mix of languages. Existing speech technologies may be inefficient in extracting information from such conversations, where the speech data is rich in diversity with multiple languages and speakers. The DISPLACE (DIarization of SPeaker and LAnguage in Conversational Environments) challenge constitutes an open-call for evaluating and bench-marking the speaker and language diarization technologies on this challenging condition. The challenge entailed two tracks: Track-1 focused on speaker diarization (SD) in multilingual situations while, Track-2 addressed the language diarization (LD) in a multi-speaker scenario. Both the tracks were evaluated using the same underlying audio data. To facilitate this evaluation, a real-world dataset featuring multilingual, multi-speaker conversational far-field speech was recorded and distributed. Furthermore, a baseline system was made available for both SD and LD task which mimicked the state-of-art in these tasks. The challenge garnered a total of $42$ world-wide registrations and received a total of $19$ combined submissions for Track-1 and Track-2. This paper describes the challenge, details of the datasets, tasks, and the baseline system. Additionally, the paper provides a concise overview of the submitted systems in both tracks, with an emphasis given to the top performing systems. The paper also presents insights and future perspectives for SD and LD tasks, focusing on the key challenges that the systems need to overcome before wide-spread commercial deployment on such conversations.

Via

Access Paper or Ask Questions

Self-Influence Guided Data Reweighting for Language Model Pre-training

Nov 02, 2023

Megh Thakkar, Tolga Bolukbasi, Sriram Ganapathy, Shikhar Vashishth, Sarath Chandar, Partha Talukdar

Figure 1 for Self-Influence Guided Data Reweighting for Language Model Pre-training

Figure 2 for Self-Influence Guided Data Reweighting for Language Model Pre-training

Figure 3 for Self-Influence Guided Data Reweighting for Language Model Pre-training

Figure 4 for Self-Influence Guided Data Reweighting for Language Model Pre-training

Abstract:Language Models (LMs) pre-trained with self-supervision on large text corpora have become the default starting point for developing models for various NLP tasks. Once the pre-training corpus has been assembled, all data samples in the corpus are treated with equal importance during LM pre-training. However, due to varying levels of relevance and quality of data, equal importance to all the data samples may not be the optimal choice. While data reweighting has been explored in the context of task-specific supervised learning and LM fine-tuning, model-driven reweighting for pre-training data has not been explored. We fill this important gap and propose PRESENCE, a method for jointly reweighting samples by leveraging self-influence (SI) scores as an indicator of sample importance and pre-training. PRESENCE promotes novelty and stability for model pre-training. Through extensive analysis spanning multiple model sizes, datasets, and tasks, we present PRESENCE as an important first step in the research direction of sample reweighting for pre-training language models.

* Accepted to EMNLP 2023

Via

Access Paper or Ask Questions

Accented Speech Recognition With Accent-specific Codebooks

Oct 27, 2023

Darshan Prabhu, Preethi Jyothi, Sriram Ganapathy, Vinit Unni

Figure 1 for Accented Speech Recognition With Accent-specific Codebooks

Figure 2 for Accented Speech Recognition With Accent-specific Codebooks

Figure 3 for Accented Speech Recognition With Accent-specific Codebooks

Figure 4 for Accented Speech Recognition With Accent-specific Codebooks

Abstract:Speech accents pose a significant challenge to state-of-the-art automatic speech recognition (ASR) systems. Degradation in performance across underrepresented accents is a severe deterrent to the inclusive adoption of ASR. In this work, we propose a novel accent adaptation approach for end-to-end ASR systems using cross-attention with a trainable set of codebooks. These learnable codebooks capture accent-specific information and are integrated within the ASR encoder layers. The model is trained on accented English speech, while the test data also contained accents which were not seen during training. On the Mozilla Common Voice multi-accented dataset, we show that our proposed approach yields significant performance gains not only on the seen English accents (up to $37\%$ relative improvement in word error rate) but also on the unseen accents (up to $5\%$ relative improvement in WER). Further, we illustrate benefits for a zero-shot transfer setup on the L2Artic dataset. We also compare the performance with other approaches based on accent adversarial training.

* Accepted to EMNLP 2023 Main Conference (Long Paper)

Via

Access Paper or Ask Questions