CoolFace
Datasetpublic

UniDataPro/real-vs-fake-human-voice-deepfake-audio

Deepfake Audio Dataset Dataset contains 5,000 audio files, comprising both authentic human recordings and synthetic** AI-generated voice** samples. It designed for advanced research in deepfake detection, focusing on detecting fake voices and generated speech analysis. Specifically engineered to challenge voice authentication systems, it supports the development of robust models for real vs fake human voice recognition. By utilizing this dataset, researchers and developers can… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/real-vs-fake-human-voice-deepfake-audio.

sourceHugging Facecc-by-nc-nd-4.0updated 1mo agoView on Hugging Face
5likes190downloads
Dataset Card

Deepfake Audio Dataset

Dataset contains 5,000 audio files, comprising both authentic human recordings and synthetic AI-generated voice samples. It designed for advanced research in deepfake detection, focusing on detecting fake voices and generated speech analysis. Specifically engineered to challenge voice authentication systems, it supports the development of robust models for real vs fake human voice recognition.

By utilizing this dataset, researchers and developers can advance the training of accurate detection models for critical applications such as mobile authentication, identity verification services, journalism verification, and monitoring social media - [Get the data](https://unidata.pro/datasets/real-vs-fake-human-voice-deepfake-audio/?utm_source=huggingface&utm_medium=referral&utm_campaign=real-vs-fake-human-voice-deepfake-audio)

Researchers can utilize this dataset to explore detection technology and recognition algorithms that aim to prevent impostor attacks and improve audio-based authentication processes.

Frequently Asked Questions

What audio formats are provided?

Deepfake audio dataset includes recordings in both M4A and MP3 formats. Supporting multiple audio formats allows researchers to evaluate whether compression and encoding influence the performance of deepfake detection models and audio forensic algorithms.

Who can benefit from this dataset?

The dataset is valuable for AI researchers, cybersecurity companies, digital forensics teams, biometric authentication providers, universities, and organizations developing voice security technologies. It supports research into deepfake detection, speaker verification, voice biometrics, and audio authenticity verification.

What metadata is included with the audio files?

Each audio file includes metadata such as country, gender, speaker ID, age, audio group, file name, and transcription text.

💵 Buy the Dataset: This is a limited preview of the data. To access the full dataset, please contact us at https://unidata.pro to discuss your requirements and pricing options.

Dataset provides a robust foundation for achieving higher detection accuracy and advancing audio deepfake detection methods, which are essential for preventing identity fraud and ensuring reliable biometric verification.

🌐 UniData provides high-quality datasets, content moderation, data collection and annotation for your AI/ML projects