Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:MTI-Net: A Multi-Target Speech Intelligibility Prediction Model

Apr 07, 2022

Ryandhimas E. Zezario, Szu-wei Fu, Fei Chen, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao

Figure 1 for MTI-Net: A Multi-Target Speech Intelligibility Prediction Model

Figure 2 for MTI-Net: A Multi-Target Speech Intelligibility Prediction Model

Figure 3 for MTI-Net: A Multi-Target Speech Intelligibility Prediction Model

Figure 4 for MTI-Net: A Multi-Target Speech Intelligibility Prediction Model

Share this with someone who'll enjoy it:

Abstract:Recently, deep learning (DL)-based non-intrusive speech assessment models have attracted great attention. Many studies report that these DL-based models yield satisfactory assessment performance and good flexibility, but their performance in unseen environments remains a challenge. Furthermore, compared to quality scores, fewer studies elaborate deep learning models to estimate intelligibility scores. This study proposes a multi-task speech intelligibility prediction model, called MTI-Net, for simultaneously predicting human and machine intelligibility measures. Specifically, given a speech utterance, MTI-Net is designed to predict subjective listening test results and word error rate (WER) scores. We also investigate several methods that can improve the prediction performance of MTI-Net. First, we compare different features (including low-level features and embeddings from self-supervised learning (SSL) models) and prediction targets of MTI-Net. Second, we explore the effect of transfer learning and multi-tasking learning on training MTI-Net. Finally, we examine the potential advantages of fine-tuning SSL embeddings. Experimental results demonstrate the effectiveness of using cross-domain features, multi-task learning, and fine-tuning SSL embeddings. Furthermore, it is confirmed that the intelligibility and WER scores predicted by MTI-Net are highly correlated with the ground-truth scores.

* Submitted to Interspeech 2022

View paper on

Share this with someone who'll enjoy it:

Title:MTI-Net: A Multi-Target Speech Intelligibility Prediction Model

Paper and Code