Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

Apr 15, 2025

Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Jianqiang Ren, Liefeng Bo, Zhigang Tu

Figure 1 for EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

Figure 2 for EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

Figure 3 for EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

Figure 4 for EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

Share this with someone who'll enjoy it:

Abstract:Masked modeling framework has shown promise in co-speech motion generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we propose a speech-queried attention-based mask modeling framework for co-speech motion generation. Our key insight is to leverage motion-aligned speech features to guide the masked motion modeling process, selectively masking rhythm-related and semantically expressive motion frames. Specifically, we first propose a motion-audio alignment module (MAM) to construct a latent motion-audio joint space. In this space, both low-level and high-level speech features are projected, enabling motion-aligned speech representation using learnable speech queries. Then, a speech-queried attention mechanism (SQA) is introduced to compute frame-level attention scores through interactions between motion keys and speech queries, guiding selective masking toward motion frames with high attention scores. Finally, the motion-aligned speech features are also injected into the generation network to facilitate co-speech motion generation. Qualitative and quantitative evaluations confirm that our method outperforms existing state-of-the-art approaches, successfully producing high-quality co-speech motion.

* 12 pages, 12 figures

View paper on

Share this with someone who'll enjoy it:

Title:EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

Paper and Code