datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Barkopedia-Dog-Vocal-Detection
🐾 Dog Vocal Detection
This dataset is curated from internet videos to support research in dog vocalization detection using both weak and strong supervision.
It contains approximately 7,500 seconds of strongly labeled training audio
Over 9,000 seconds of weakly labeled clips sourced from AudioSet are included.
The dataset also provides 24 hours of unlabeled audio clips from our own collection.
To simulate realistic conditions, some clips feature dogs present without barking… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia-Dog-Vocal-Detection.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic voice… See the full description on the dataset page: https://huggingface.co/datasets/garystafford/deepfake-audio-detection.drone-audio-detection-samples
Dataset Description
Drone Audio Detection Samples (DADS) is currently the largest publicly available drone audio database, specifically designed for developing drone detection systems using deep learning techniques. All audio files are standardized to a sample rate of 16,000 Hz, 16-bit depth, mono-channel, and vary in length from 500 milliseconds to several minutes.
Most drone audio files were manually trimmed to ensure that a drone was always present in the recording. However, some… See the full description on the dataset page: https://huggingface.co/datasets/geronimobasso/drone-audio-detection-samples.deepfake_detection_dataset_urdu
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset
This repository contains the Urdu Deepfake Audio Dataset introduced in the ACL 2024 paper "Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset".
The dataset focuses on two spoofing attacks – Tacotron and VITS TTS – and includes bonafide audio samples for comparison. The dataset construction ensures phonemic cover and balance, making it suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/CSALT/deepfake_detection_dataset_urdu.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1LMD-AI-Detection
LMD AI-Generated Music Detection Benchmark
(Note: The corresponding research paper will be released later.)
Dataset Description
The rapid advancement of AI music generation has raised growing concerns about the authenticity of digital music. While deepfake detection has been extensively studied in the audio domain, symbolic music (MIDI) remains largely unexplored.
This dataset presents a comprehensive benchmark for AI-generated symbolic music detection, examining… See the full description on the dataset page: https://huggingface.co/datasets/dhlee3000/LMD-AI-Detection.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic… See the full description on the dataset page: https://huggingface.co/datasets/koyyalamudiraghavendra/deepfake-audio-detection.arabic_commands_detection
Dataset Card for "arabic_commands_detection"
More Information needed
mixed-language-detection-pilot-complete-sentences
Mixed-Language Speech Detection Pilot — Complete Sentences
This is the complete-sentence revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed
Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.mixed-language-detection-pilot-fleurs-voices
Mixed-Language Speech Detection Pilot — Native FLEURS Voices
This is the native-reference revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed in this revision
Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.mixed-language-detection-english-accented-vc
Mixed-Language Speech Detection Pilot
This dataset is a 6,000-clip binary audio-classification pilot for detecting
whether an utterance contains one language (label = 0) or more than one
language (label = 1). It covers Turkish (tur), Northern Kurdish/Kurmanji
(kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and
English (eng).
Dataset composition
Construction
Mixed
Monolingual
Total
Single-call OmniVoice
500
500
1,000
Segment-level… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-english-accented-vc.Audio-based-Voilence-detection-DatasetDataset_audio_threat_detectionvessel-detection-datasetaudio-emotion-detection-dataset
Audio Emotion Detection Dataset
Github: Audio Emotion Detection Dataset
Connect with me : Linkedin
Speech clips in English and Hindi annotated with emotion labels and ASR transcripts.
Audio is sourced from public YouTube videos and trimmed to approximately 60 seconds per clip.
Noise reduction is applied via noisereduce and silero-vad.
Emotions (5 classes)
Label
Description
angry
Aggressive, confrontational speech
calm… See the full description on the dataset page: https://huggingface.co/datasets/RapidOrc121/audio-emotion-detection-dataset.Key_Detectioninfant-cry-detection-dataset
Infant Cry Detection Dataset — 50+ Hours of Real Baby Cry Audio
50+ hours of real infant cry recordings for training infant cry detection, cry classification, and sound event detection models. Manually verified files captured in natural domestic conditions, with per-file metadata on location, background noise, and recording device
Contact us and share your feedback — receive additional samples for free! 😊
Key Highlights
50+ hours of real-world… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/infant-cry-detection-dataset.turn-end-detection
Turn-end detection from real ASR prefixes, with audio
100,348 labelled end-of-turn decision points over
53,140 synthesized customer-service utterances, each one paired
with the 16 kHz audio it was cut from, plus the endpointing decisions
6 commercial endpointer configurations
made on the same audio.
The question each row poses is the one a voice agent has to answer continuously:
given everything heard so far, has the caller finished speaking? Ending the
turn too early talks over… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/turn-end-detection.fake-audio-detection-augmented2realtime-turn-detection-test-data
Realtime speech test recordings
Synthetic speech recordings for black-box Realtime API behavior tests in
Speaches. Each WAV file is the unmodified output of OpenAI
text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk
boundaries for their scenarios.
metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file
digest, expected text, transcription, and word/speech intervals from… See the full description on the dataset page: https://huggingface.co/datasets/speaches-ai/realtime-turn-detection-test-data.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Tanishq125/deepfake-audio-detection.scream_detection_heavy_metal
Dataset card for Scream Detection in Heavy Metal Music
This dataset contains the processed dataset used in the paper "Scream Detection in Heavy Metal Music" (Kalbag & Lerch, 2022) from the Georgia Institute of Technology.
This dataset contains annotations of 57 songs, distributed over 34 bands and 47 albums. The vocal events are labelled into 5 classes:
Clean (or sung vocal)
Low Fry Scream
Mid Fry Scream
High Fry Scream
Layered Vocals
The label "Layered Vocals" has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/jpdiazpardo/scream_detection_heavy_metal.clap-detectionVocal_Technique_Detectionaudio-emotion-detection-dataset
Audio Emotion Detection Dataset
Github: Audio Emotion Detection Dataset
Connect with me : Linkedin
Speech clips in English and Hindi annotated with emotion labels and ASR transcripts.
Audio is sourced from public YouTube videos and trimmed to approximately 60 seconds per clip.
Noise reduction is applied via noisereduce and silero-vad.
Emotions (5 classes)
Label
Description
angry
Aggressive, confrontational speech
calm… See the full description on the dataset page: https://huggingface.co/datasets/kotangalechaitali9007/audio-emotion-detection-dataset.urdu-turn-detection-audio-v2
🗣️ Urdu Turn Detection (Audio Dataset V2)
This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech.
It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection.
🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.drone-audio-detection-samples
Dataset Description
Drone Audio Detection Samples (DADS) is currently the largest publicly available drone audio database, specifically designed for developing drone detection systems using deep learning techniques. All audio files are standardized to a sample rate of 16,000 Hz, 16-bit depth, mono-channel, and vary in length from 500 milliseconds to several minutes.
Most drone audio files were manually trimmed to ensure that a drone was always present in the recording. However… See the full description on the dataset page: https://huggingface.co/datasets/gr1047/drone-audio-detection-samples.Speech_Concurrent_Event_Detectionspeech-deepfake-detection-40kVocal_Fry_Detection
