Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Brendan Fahy

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Jun 13, 2025

Haoran Zhou, Xingchen Song, Brendan Fahy, Qiaochu Song, Binbin Zhang, Zhendong Peng, Anshul Wadhawan, Denglin Jiang, Apurv Verma, Vinay Ramesh(+2 more)

Figure 1 for Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Figure 2 for Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Figure 3 for Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Figure 4 for Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Abstract:OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence-to-sequence objective, lacks native support for streaming ASR. In this paper, we fine-tune Whisper for streaming ASR using the WeNet toolkit by adopting a Unified Two-pass (U2) structure. We introduce an additional Connectionist Temporal Classification (CTC) decoder trained with causal attention masks to generate streaming partial transcripts, while the original Whisper decoder reranks these partial outputs. Our experiments on LibriSpeech and an earnings call dataset demonstrate that, with adequate fine-tuning data, Whisper can be adapted into a capable streaming ASR model. We also introduce a hybrid tokenizer approach, which uses a smaller token space for the CTC decoder while retaining Whisper's original token space for the attention decoder, resulting in improved data efficiency and generalization.

* Accepted to INTERSPEECH 2025

Via

Access Paper or Ask Questions

Dialogue Act Classification in Group Chats with DAG-LSTMs

Aug 02, 2019

Ozan İrsoy, Rakesh Gosangi, Haimin Zhang, Mu-Hsin Wei, Peter Lund, Duccio Pappadopulo, Brendan Fahy, Neophytos Nephytou, Camilo Ortiz

Figure 1 for Dialogue Act Classification in Group Chats with DAG-LSTMs

Figure 2 for Dialogue Act Classification in Group Chats with DAG-LSTMs

Figure 3 for Dialogue Act Classification in Group Chats with DAG-LSTMs

Figure 4 for Dialogue Act Classification in Group Chats with DAG-LSTMs

Abstract:Dialogue act (DA) classification has been studied for the past two decades and has several key applications such as workflow automation and conversation analytics. Researchers have used, to address this problem, various traditional machine learning models, and more recently deep neural network models such as hierarchical convolutional neural networks (CNNs) and long short-term memory (LSTM) networks. In this paper, we introduce a new model architecture, directed-acyclic-graph LSTM (DAG-LSTM) for DA classification. A DAG-LSTM exploits the turn-taking structure naturally present in a multi-party conversation, and encodes this relation in its model structure. Using the STAC corpus, we show that the proposed method performs roughly 0.8% better in accuracy and 1.2% better in macro-F1 score when compared to existing methods. The proposed method is generic and not limited to conversation applications.

* Appeared in SIGIR 2019 Workshop on Conversational Interaction Systems

Via

Access Paper or Ask Questions