Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis

Dec 09, 2020

Anurag Chowdhury, Arun Ross, Prabu David

Figure 1 for DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis

Figure 2 for DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis

Figure 3 for DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis

Figure 4 for DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis

Share this with someone who'll enjoy it:

Abstract:Automatic speaker recognition algorithms typically use physiological speech characteristics encoded in the short term spectral features for characterizing speech audio. Such algorithms do not capitalize on the complementary and discriminative speaker-dependent characteristics present in the behavioral speech features. In this work, we propose a prosody encoding network called DeepTalk for extracting vocal style features directly from raw audio data. The DeepTalk method outperforms several state-of-the-art physiological speech characteristics-based speaker recognition systems across multiple challenging datasets. The speaker recognition performance is further improved by combining DeepTalk with a state-of-the-art physiological speech feature-based speaker recognition system. We also integrate the DeepTalk method into a current state-of-the-art speech synthesizer to generate synthetic speech. A detailed analysis of the synthetic speech shows that the DeepTalk captures F0 contours essential for vocal style modeling. Furthermore, DeepTalk-based synthetic speech is shown to be almost indistinguishable from real speech in the context of speaker recognition.

* IEEE ICASSP 2020 Submission, 5 pages, 3 figures

View paper on

Share this with someone who'll enjoy it:

Title:DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis

Paper and Code