Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Peizhuo Liu

Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

May 21, 2025

Yuhao Zhang, Xiangnan Ma, Kaiqi Kou, Peizhuo Liu, Weiqiao Shan, Benyou Wang, Tong Xiao, Yuxin Huang, Zhengtao Yu, Jingbo Zhu

Figure 1 for Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

Figure 2 for Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

Figure 3 for Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

Figure 4 for Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

Abstract:The success of building textless speech-to-speech translation (S2ST) models has attracted much attention. However, S2ST still faces two main challenges: 1) extracting linguistic features for various speech signals, called cross-modal (CM), and 2) learning alignment of difference languages in long sequences, called cross-lingual (CL). We propose the unit language to overcome the two modeling challenges. The unit language can be considered a text-like representation format, constructed using $n$-gram language modeling. We implement multi-task learning to utilize the unit language in guiding the speech modeling process. Our initial results reveal a conflict when applying source and target unit languages simultaneously. We propose task prompt modeling to mitigate this conflict. We conduct experiments on four languages of the Voxpupil dataset. Our method demonstrates significant improvements over a strong baseline and achieves performance comparable to models trained with text.

* Accepted to ACL 2025 Findings

Via

Access Paper or Ask Questions

SpMis: An Investigation of Synthetic Spoken Misinformation Detection

Sep 17, 2024

Peizhuo Liu, Li Wang, Renqiang He, Haorui He, Lei Wang, Huadi Zheng, Jie Shi, Tong Xiao, Zhizheng Wu

Figure 1 for SpMis: An Investigation of Synthetic Spoken Misinformation Detection

Figure 2 for SpMis: An Investigation of Synthetic Spoken Misinformation Detection

Figure 3 for SpMis: An Investigation of Synthetic Spoken Misinformation Detection

Figure 4 for SpMis: An Investigation of Synthetic Spoken Misinformation Detection

Abstract:In recent years, speech generation technology has advanced rapidly, fueled by generative models and large-scale training techniques. While these developments have enabled the production of high-quality synthetic speech, they have also raised concerns about the misuse of this technology, particularly for generating synthetic misinformation. Current research primarily focuses on distinguishing machine-generated speech from human-produced speech, but the more urgent challenge is detecting misinformation within spoken content. This task requires a thorough analysis of factors such as speaker identity, topic, and synthesis. To address this need, we conduct an initial investigation into synthetic spoken misinformation detection by introducing an open-source dataset, SpMis. SpMis includes speech synthesized from over 1,000 speakers across five common topics, utilizing state-of-the-art text-to-speech systems. Although our results show promising detection capabilities, they also reveal substantial challenges for practical implementation, underscoring the importance of ongoing research in this critical area.

* Accepted in SLT 2024

Via

Access Paper or Ask Questions