Picture for Keqi Deng

Keqi Deng

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

Add code
Jul 20, 2026
Viaarxiv icon

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

Add code
Jul 07, 2026
Viaarxiv icon

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

Add code
Jul 02, 2026
Viaarxiv icon

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

Add code
May 29, 2026
Viaarxiv icon

UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training

Add code
May 26, 2026
Viaarxiv icon

Speech LLMs are Contextual Reasoning Transcribers

Add code
Apr 01, 2026
Viaarxiv icon

SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation

Add code
Apr 22, 2025
Figure 1 for SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
Figure 2 for SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
Figure 3 for SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
Figure 4 for SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
Viaarxiv icon

Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition

Add code
Dec 21, 2024
Figure 1 for Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
Figure 2 for Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
Figure 3 for Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
Figure 4 for Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
Viaarxiv icon

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Add code
Oct 09, 2024
Figure 1 for F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Figure 2 for F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Figure 3 for F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Figure 4 for F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Viaarxiv icon

CoT-ST: Enhancing LLM-based Speech Translation with Multimodal Chain-of-Thought

Add code
Sep 29, 2024
Figure 1 for CoT-ST: Enhancing LLM-based Speech Translation with Multimodal Chain-of-Thought
Figure 2 for CoT-ST: Enhancing LLM-based Speech Translation with Multimodal Chain-of-Thought
Figure 3 for CoT-ST: Enhancing LLM-based Speech Translation with Multimodal Chain-of-Thought
Figure 4 for CoT-ST: Enhancing LLM-based Speech Translation with Multimodal Chain-of-Thought
Viaarxiv icon