AUDETER: A large-scale dataset for deepfake audio detection in open worlds
Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
8-2026
Abstract
Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving distribution shifts between training and test data, driven by the complexity of human speech and the rapid evolution of synthesis systems. Existing datasets suffer from limited real speech diversity, insufficient coverage of recent synthesis systems, and heterogeneous mixtures of deepfake sources, which hinder systematic evaluation and open-world model training. To address these issues, we introduce AUDETER (AUdio DEepfake TEst Range), a large-scale and highly diverse deepfake audio dataset comprising over 4,500 hours of synthetic audio generated by 11 recent TTS models and 10 vocoders, totalling 3 million clips. We further observe that most existing detectors default to binary supervised training, which can induce negative transfer across synthesis sources when the training data contains highly diverse deepfake patterns, impacting overall generalisation. As a complementary contribution, we propose an effective curriculum-learning-based approach to mitigate this effect. Extensive experiments show that existing detection models struggle to generalise to novel deepfakes and human speech in AUDETER, whereas XLR-based detectors trained on AUDETER achieve strong cross-domain performance across multiple benchmarks, achieving an EER of 1.87% on In-the-Wild. AUDETER is available on GitHub: https://github.com/mike-qz-wang/AUDETER.
Keywords
deepfake audio detection, open-world detection, large-scale dataset
Discipline
Artificial Intelligence and Robotics | Information Security
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
KDD '26: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, Jeju Island, South Korea, August 9-13
First Page
9939
Last Page
9950
Identifier
10.1145/3770855.3817583
Publisher
ACM
City or Country
New York
Citation
WANG, Qizhou; HUANG, Hanxun; PANG, Guansong; ERFANI, Sarah; and LECKIE, Christopher.
AUDETER: A large-scale dataset for deepfake audio detection in open worlds. (2026). KDD '26: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, Jeju Island, South Korea, August 9-13. 9939-9950.
Available at: https://ink.library.smu.edu.sg/sis_research/11300
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1145/3770855.3817583