speech recognition


Speech recognition is the task of identifying words spoken aloud, analyzing the voice and language, and accurately transcribing the words.

Using Songs to Improve Kazakh Automatic Speech Recognition

Add code
Mar 03, 2026
Viaarxiv icon

Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge

Add code
Mar 05, 2026
Viaarxiv icon

Visual-Informed Speech Enhancement Using Attention-Based Beamforming

Add code
Mar 05, 2026
Viaarxiv icon

When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper

Add code
Mar 05, 2026
Viaarxiv icon

PersianPunc: A Large-Scale Dataset and BERT-Based Approach for Persian Punctuation Restoration

Add code
Mar 05, 2026
Viaarxiv icon

WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech

Add code
Mar 05, 2026
Viaarxiv icon

Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach

Add code
Mar 13, 2026
Viaarxiv icon

BabAR: from phoneme recognition to developmental measures of young children's speech production

Add code
Mar 05, 2026
Viaarxiv icon

Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards

Add code
Mar 05, 2026
Viaarxiv icon

An Investigation Into Various Approaches For Bengali Long-Form Speech Transcription and Bengali Speaker Diarization

Add code
Mar 03, 2026
Viaarxiv icon