Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Unmasking real-world audio deepfakes: A data-centric approach

Jun 11, 2025

David Combei, Adriana Stan, Dan Oneata, Nicolas Müller, Horia Cucu

Figure 1 for Unmasking real-world audio deepfakes: A data-centric approach

Figure 2 for Unmasking real-world audio deepfakes: A data-centric approach

Figure 3 for Unmasking real-world audio deepfakes: A data-centric approach

Figure 4 for Unmasking real-world audio deepfakes: A data-centric approach

Share this with someone who'll enjoy it:

Abstract:The growing prevalence of real-world deepfakes presents a critical challenge for existing detection systems, which are often evaluated on datasets collected just for scientific purposes. To address this gap, we introduce a novel dataset of real-world audio deepfakes. Our analysis reveals that these real-world examples pose significant challenges, even for the most performant detection models. Rather than increasing model complexity or exhaustively search for a better alternative, in this work we focus on a data-centric paradigm, employing strategies like dataset curation, pruning, and augmentation to improve model robustness and generalization. Through these methods, we achieve a 55% relative reduction in EER on the In-the-Wild dataset, reaching an absolute EER of 1.7%, and a 63% reduction on our newly proposed real-world deepfakes dataset, AI4T. These results highlight the transformative potential of data-centric approaches in enhancing deepfake detection for real-world applications. Code and data available at: https://github.com/davidcombei/AI4T.

* Accepted at Interspeech 2025

View paper on

Share this with someone who'll enjoy it:

Title:Unmasking real-world audio deepfakes: A data-centric approach

Paper and Code