Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Jio Gim

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Aug 09, 2025

Jinsung Yoon, Wooyeol Jeong, Jio Gim, Young-Joo Suh

Figure 1 for Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Figure 2 for Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Figure 3 for Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Figure 4 for Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Abstract:Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotional style using distinct references, is crucial. However, existing methods often struggle to fully disentangle these attributes and lack the ability to model fine-grained emotional expressions such as temporal dynamics. We propose Maestro-EVC, a controllable EVC framework that enables independent control of content, speaker identity, and emotion by effectively disentangling each attribute from separate references. We further introduce a temporal emotion representation and an explicit prosody modeling with prosody augmentation to robustly capture and transfer the temporal dynamics of the target emotion, even under prosody-mismatched conditions. Experimental results confirm that Maestro-EVC achieves high-quality, controllable, and emotionally expressive speech synthesis.

* Accepted at ASRU 2025

Via

Access Paper or Ask Questions

AutoCycle-VC: Towards Bottleneck-Independent Zero-Shot Cross-Lingual Voice Conversion

Oct 10, 2023

Haeyun Choi, Jio Gim, Yuho Lee, Youngin Kim, Young-Joo Suh

Figure 1 for AutoCycle-VC: Towards Bottleneck-Independent Zero-Shot Cross-Lingual Voice Conversion

Figure 2 for AutoCycle-VC: Towards Bottleneck-Independent Zero-Shot Cross-Lingual Voice Conversion

Figure 3 for AutoCycle-VC: Towards Bottleneck-Independent Zero-Shot Cross-Lingual Voice Conversion

Figure 4 for AutoCycle-VC: Towards Bottleneck-Independent Zero-Shot Cross-Lingual Voice Conversion

Abstract:This paper proposes a simple and robust zero-shot voice conversion system with a cycle structure and mel-spectrogram pre-processing. Previous works suffer from information loss and poor synthesis quality due to their reliance on a carefully designed bottleneck structure. Moreover, models relying solely on self-reconstruction loss struggled with reproducing different speakers' voices. To address these issues, we suggested a cycle-consistency loss that considers conversion back and forth between target and source speakers. Additionally, stacked random-shuffled mel-spectrograms and a label smoothing method are utilized during speaker encoder training to extract a time-independent global speaker representation from speech, which is the key to a zero-shot conversion. Our model outperforms existing state-of-the-art results in both subjective and objective evaluations. Furthermore, it facilitates cross-lingual voice conversions and enhances the quality of synthesized speech.

Via

Access Paper or Ask Questions