Picture for Shansong Liu

Shansong Liu

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

Add code
Aug 10, 2026
Viaarxiv icon

Unified Audio Generation and Editing via Joint Condition Modeling and Progressive Training

Add code
Jun 15, 2026
Viaarxiv icon

High-Fidelity Generative Audio Compression at 0.275kbps

Add code
Jan 31, 2026
Viaarxiv icon

Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models

Add code
Dec 26, 2025
Viaarxiv icon

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Add code
Mar 11, 2025
Viaarxiv icon

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Add code
Dec 09, 2024
Viaarxiv icon

Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer

Add code
Oct 07, 2024
Figure 1 for Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
Figure 2 for Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
Figure 3 for Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
Viaarxiv icon

M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Add code
Nov 28, 2023
Figure 1 for M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
Figure 2 for M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
Figure 3 for M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
Figure 4 for M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
Viaarxiv icon

HumTrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond

Add code
Sep 18, 2023
Figure 1 for HumTrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond
Figure 2 for HumTrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond
Figure 3 for HumTrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond
Figure 4 for HumTrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond
Viaarxiv icon

Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

Add code
Aug 22, 2023
Viaarxiv icon