datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
excavationpro-music-stream
Excavationpro public music stream (160 kbps)
Owner / artist: Justin Helmer · Excavationpro · LightfatherPolicy: Own-work only. Public discovery streams (not DistroKid-dependent).Lattice signature: Δ9Φ963-PUBLIC-MUSIC-STREAM-v1
Listen
https://deepseekoracle.github.io/Excavationpro/excavationpro-listen.html
http://asiancoastline.com/ (custom domain music portal)
Layout
Path
Role
stream/<sha256>.mp3
Flat 160k streams (~first 10k −… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/excavationpro-music-stream.DeepASMR-datasetDeepDialogue-orpheus
DeepDialogue-orpheus
DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text.
🚨 Important Notice
This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.himalaya-ai-stt-datasethimalaya-ai-final-sttmtedx
Multilingual TEDx (mTEDx) — SLR100
mTEDx is a multilingual speech recognition and translation corpus built from
TEDx Talks.Original resource: https://www.openslr.org/100/
The corpus provides audio recordings and VTT transcripts for 8 languages
(Spanish, French, Portuguese, Italian, Russian, Greek, Arabic, German) with
aligned translations into up to 5 languages (English, Spanish, French,
Portuguese, Italian).
License: CC BY-NC-ND 4.0Contact: Elizabeth Salesky (esalesky@jhu.edu)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/mtedx.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic voice… See the full description on the dataset page: https://huggingface.co/datasets/garystafford/deepfake-audio-detection.lygo-protocol-stack
LYGO Protocol Stack — Hugging Face mirror
Canonical GitHub: DeepSeekOracle/lygo-protocol-stack
Contents
P0 Nano Kernel — Φ-gate (AMPLIFY / SOFTEN / QUARANTINE), 42 hardened test vectors, Python/C/Rust parity
P1–P5 — Memory Mycelium, Cognitive Bridge, Vortex Consensus, Ascension Engine, Harmony Node
ClawHub catalog — links + skills.json (full skill trees on GitHub)
P0 quick demo (local)
pip install pytest
python tools/run_p0_demo.py
python… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/lygo-protocol-stack.IndicTTS-Deepfake-Challenge-Data
IndicTTS Deepfake Detection Challenge
Participants will use the SherryT997/IndicTTS-Deepfake-Challenge-Data dataset, hosted on Hugging Face. This dataset consists of train and test splits and contains speech samples in 16 Indian languages, along with metadata for each audio clip.
🚀 Dataset to Use: SherryT997/IndicTTS-Deepfake-Challenge-Data
This is the official dataset for the challenge and must be used for training and evaluation.
📌 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SherryT997/IndicTTS-Deepfake-Challenge-Data.BRSpeech-DF
🗣️ BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese
🧩 Description
BRSpeech-DF is the first publicly available dataset for deepfake speech detection in Portuguese, covering both Brazilian and European variants.
It contains 459,000 audio samples, including both real and synthetic speech generated using multiple zero-shot text-to-speech (TTS) models.
This dataset aims to foster the development of more robust, inclusive, and multilingual audio deepfake… See the full description on the dataset page: https://huggingface.co/datasets/AKCIT-Deepfake/BRSpeech-DF.DeepVoice
DeepVoice
Benchmark-ready packaging of the DEEP-VOICE real-vs-AI-generated speech dataset
(arXiv 2308.12734), for speech anti-spoofing and
synthetic / deepfake voice detection.
Overview
DEEP-VOICE is a binary-classification benchmark: bonafide (genuine human speech)
vs. spoof (AI voice-converted speech). The spoof side is generated with
Retrieval-based Voice Conversion (RVC), converting one real speaker's recording
into the voice of another; the bonafide side is… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/DeepVoice.deepfake_detection_dataset_urdu
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset
This repository contains the Urdu Deepfake Audio Dataset introduced in the ACL 2024 paper "Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset".
The dataset focuses on two spoofing attacks – Tacotron and VITS TTS – and includes bonafide audio samples for comparison. The dataset construction ensures phonemic cover and balance, making it suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/CSALT/deepfake_detection_dataset_urdu.NonverbalTTS
NonverbalTTS Dataset 🎵🗣️
NonverbalTTS is a 17-hour open-access English speech corpus with aligned text annotations for nonverbal vocalizations (NVs) and emotional categories, designed to advance expressive text-to-speech (TTS) research.
Key Features ✨
17 hours of high-quality speech data
10 NV types: Breathing, laughter, sighing, sneezing, coughing, throat clearing, groaning, grunting, snoring, sniffing8 emotion categories: Angry, disgusted, fearful, happy, neutral… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/NonverbalTTS.mls_it_pseudo_labelled-large-v3DeepASMR-dataset-backuparknights_voices_zh
ZH Voice-Text Dataset for Arknights Waifus
This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
12431 records, 25.9 hours in total. Average duration is approximately 7.49s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.nepali-audio-deepfake-datasetcommon_voice_26_0
Dataset Card for Common Voice Corpus 26.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 26. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/common_voice_26_0.microsoft-speech-corpus-indian
Microsoft Speech Corpus – Indian Languages
Dataset Description
This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript.
Attribution required: "Data provided by Microsoft and SpeechOcean.com"
⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.Deepfake-Eval-2024
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
Deepfake-Eval-2024 is an in-the-wild deepfake dataset. Deepfake-Eval-2024 contains 44 hours of videos, 56.5 hours of audio, and 1,975 images, encompassing contemporary manipulation technologies, diverse media content, 88 different website sources, and 52 different languages. Deepfake-Eval-2024 contains manually labeled real and fake media. Deepfake-Eval-2024 is designed to facilitate deepfake… See the full description on the dataset page: https://huggingface.co/datasets/nuriachandra/Deepfake-Eval-2024.real-vs-fake-human-voice-deepfake-audio
Deepfake Audio Dataset
Dataset contains 5,000 audio files, comprising both authentic human recordings and synthetic** AI-generated voice** samples. It designed for advanced research in deepfake detection, focusing on detecting fake voices and generated speech analysis. Specifically engineered to challenge voice authentication systems, it supports the development of robust models for real vs fake human voice recognition.
By utilizing this dataset, researchers and developers can… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/real-vs-fake-human-voice-deepfake-audio.iisc-mile-tamil-asr
IISc-MILE Tamil ASR Corpus
Summary
The IISc-MILE Tamil ASR Corpus is a large-scale transcribed speech dataset designed for training Automatic Speech Recognition (ASR) systems for the Tamil language. It contains approximately 150 hours of high-quality read speech collected from 531 speakers in a noise-free recording environment using professional USB microphones.
This corpus was published by the Medical Intelligence and Language Engineering (MILE) Lab at the Indian… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/iisc-mile-tamil-asr.azurlane_voices_jp
JP Voice-Text Dataset for Azur Lane Waifus
This is the JP voice-text dataset for azur lane playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
30160 records, 75.8 hours in total. Average duration is approximately 9.05s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/azurlane_voices_jp.Wikipedia-FA-EN-DeepSeek-V4-Flash-0731
Wikipedia Persian to English — DeepSeek V4 Flash 0731
Rolling, machine-generated English translations of Persian Wikipedia articles
from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration
full_articles_fa_without_en. 129,816 translations are
currently published in 26 immutable Parquet shards.
The target release contains 129,816 translations;
five source rows have empty plain_text and are not translated. Shards are
published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic… See the full description on the dataset page: https://huggingface.co/datasets/koyyalamudiraghavendra/deepfake-audio-detection.DeepDialogue-xttsclotho-momentcommon_voice_17_0_pseudo_labelled-large-v3DeepDialogue-xtts
DeepDialogue-xtts
DeepDialogue-xtts is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions.
This repository contains the XTTS-v2 variant of the dataset, where speech is generated using XTTS-v2 with explicit emotional conditioning.
🚨 Important
This dataset is large (~180GB) due to the inclusion of high-quality audio files. When cloning the… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-xtts.Deepfake-Eval-2024-audio
