Abstract:Audio recorded during real-time magnetic resonance imaging (rtMRI) is heavily contaminated by scanner noise, but it remains unclear whether general-purpose speech enhancement improves the signal for speech research and downstream processing. Three off-the-shelf systems---Denoiser, PASE, and RE-USE---are evaluated across five rtMRI corpora using naturally recorded inputs, a clean-input probe, and an archived paired additive-noise probe. The multi-task evaluation spans learned quality predictors, speaker and phone representations, reference-based intelligibility and quality measures, acoustic--phonetic probes, automatic speech recognition (ASR), and paralinguistic tasks. The central result is that enhancement effects are endpoint dependent: higher predicted-quality scores do not reliably imply better ASR performance or greater source fidelity. Across 15 corpus--recognizer comparisons using corpus-provided processed inputs, RE-USE yielded lower word-error-rate point estimates in 11, whereas Denoiser yielded higher estimates in 13. In the paired additive-noise probe, PASE and RE-USE improved recognized-phone agreement, intelligibility, and perceptual-quality point estimates. Denoiser improved recognized-phone agreement and short-time objective intelligibility (STOI) but reduced speaker-embedding similarity. No system was uniformly best across corpora, recognizers, and endpoints. Enhanced rtMRI audio should therefore be treated as a task-specific transformed derivative rather than a universally improved replacement for the original or DSP-processed waveform.
Abstract:Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the acoustic speech signal is the most accessible product of the speech production act, it does not directly reveal its causal neurophysiological substrates. We present the first simultaneous acquisition of real-time (dynamic) MRI, EEG, and surface EMG, capturing several key aspects of the speech production chain: brain signals, muscle activations, and articulatory movements. This multimodal acquisition paradigm presents substantial technical challenges, including MRI-induced electromagnetic interference and myogenic artifacts. To mitigate these, we introduce an artifact suppression pipeline tailored to this tri-modal setting. Once fully developed, this framework is poised to offer an unprecedented window into speech neuroscience and insights leading to brain-computer interface advances.