datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodecFake
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
Paper,
Code,
Project Page
Interspeech 2024
TL;DR: We show that better detection of deepfake speech from codec-based TTS systems can be achieved by training models on speech re-synthesized with neural audio codecs.
This dataset is released for this purpose.
See our paper and Github for more details on using our dataset.
Acknowledgement… See the full description on the dataset page: https://huggingface.co/datasets/rogertseng/CodecFake.codecfake-audio
Codecfake Dataset
Overview
The Codecfake dataset is a large-scale dataset designed for the detection of Audio Language Model (ALM)-based deepfake audio. This dataset includes millions of audio samples across two languages and various test conditions, tailored specifically for ALM-based audio detection.
Conversion
The original dataset was downloaded from Zenodo and converted to FLAC format to maintain audio quality while reducing file size. The dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/ajaykarthick/codecfake-audio.CodecFake_Plus_Dataset
Codecfake+ Dataset
Overview
This is the official dataset repository for CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset.It stores and provides access to the dataset (CoRS and CoSG), including audio samples and accompanying protocol/label files.
News
[2025.10] — CoRS and CoSG dataset and corresponding label files have been uploaded.
[2025.09] — Released public audio samples of the CoRS subset.
Download
We provide two… See the full description on the dataset page: https://huggingface.co/datasets/CodecFake/CodecFake_Plus_Dataset.CodecFake_wavs
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
Paper,
Code,
Project Page
Interspeech 2024
TL;DR: We show that better detection of deepfake speech from codec-based TTS systems can be achieved by training models on speech re-synthesized with neural audio codecs.
This dataset is released for this purpose.
See our paper and Github for more details on using our dataset.
Acknowledgement… See the full description on the dataset page: https://huggingface.co/datasets/rogertseng/CodecFake_wavs.CodecFakeCodecfakeExpressive_CodecFake
Expressive CodecFake
Expressive CodecFake is a codec-fake expressive speech dataset for audio deepfake detection research. The dataset contains codec-generated expressive speech and nonverbal vocalization samples organized into verified TAR shards.
Current Dataset Structure
Expressive_CodecFake/
├── Verbal speech CF/
│ ├── emodb_2.0_CF/
│ │ ├── emodb_2.0_CF-0000.tar
│ │ └── ...
│ ├── EMOVO_CF/
│ │ ├── EMOVO_CF-0000.tar
│ │ └── ...
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/ggirishg/Expressive_CodecFake.Union-Codecfake-DatasetSEA-Codecfake-Dataset
