Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Self-supervised Learning of Visual Speech Features with Audiovisual Speech Enhancement

May 06, 2020

Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald, Erik Marchi, Sachin Kajarekar, Devang Naik, Ahmed Hussen Abdelaziz

Figure 1 for Self-supervised Learning of Visual Speech Features with Audiovisual Speech Enhancement

Figure 2 for Self-supervised Learning of Visual Speech Features with Audiovisual Speech Enhancement

Figure 3 for Self-supervised Learning of Visual Speech Features with Audiovisual Speech Enhancement

Figure 4 for Self-supervised Learning of Visual Speech Features with Audiovisual Speech Enhancement

Share this with someone who'll enjoy it:

Abstract:We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show that visual features provide not only high-level information about speech activity, i.e. speech vs. no speech, but also fine-grained visual information about the place of articulation. An interesting byproduct of this finding is that the learned visual embeddings can be used as features for other visual speech applications. We demonstrate the effectiveness of the learned visual representations for classifying visemes (the visual analogy to phonemes). Our results provide insight into important aspects of audiovisual speech enhancement and demonstrate how such models can be used for self-supervision tasks for visual speech applications.

* Submitted to INTERSPEECH

View paper on

Share this with someone who'll enjoy it:

Title:Self-supervised Learning of Visual Speech Features with Audiovisual Speech Enhancement

Paper and Code