Picture for Francesco Verdini

Francesco Verdini

GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

Add code
Jul 02, 2026
Viaarxiv icon

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

Add code
Jun 26, 2026
Viaarxiv icon

How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not

Add code
Sep 25, 2024
Figure 1 for How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
Figure 2 for How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
Figure 3 for How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
Figure 4 for How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
Viaarxiv icon