datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.Nemotron-Content-Safety-Audio-Dataset
Nemotron Content Safety Audio Dataset
Dataset Description
The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories.
LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.serena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.CREMA-D
CREMA-D
This is an audio classification dataset for Emotion Recognition.
Classes = 6 , Split = Train-Test
Structure
audios folder contains audio files.
train.csv for training split and test.csv for the testing split.
Download
import os
import huggingface_hub
audio_datasets_path = "DATASET_PATH/Audio-Datasets"
if not os.path.exists(audio_datasets_path): print(f"Given {audio_datasets_path=} does not exist. Specify a valid path ending with 'Audio-Datasets'… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/CREMA-D.vox-cloned-data
CommonVoice Clones
This dataset consists of recordings taken from the CommonVoice english dataset.
Each voice and transcript are used as input to a voice cloner, and generate a cloned version of the voice and text.
TTS Models
We use the following high-scoring models from the TTS leaderboard:
playHT
metavoice
StyleTTSv2
XttsV2
Model Comparisons
To facilitate data exploration, check out this HF space 🤗, which allows you to listen to all clones from a given… See the full description on the dataset page: https://huggingface.co/datasets/jerpint/vox-cloned-data.central-kurdish-tts4all
TTS4All Central Kurdish Speech Dataset
Dataset Summary
The TTS4All Central Kurdish Speech Dataset is a multi-speaker speech corpus developed for speech synthesis and speech technology research in Central Kurdish (Sorani Kurdish).
The dataset was created within the TTS4All initiative during the JSALT 2025 Workshop and provides more than 35 hours of transcribed speech from three native Central Kurdish speakers.
The corpus was designed to support:
Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-tts4all.friend-bench
Can a model — or a human — tell how two people are related from a 20-second clip of how they interact?
🌐 Built on Seamless Interaction
FriendBench is a suite of benchmarks for social perception from thin-slice dyadic
interaction — inferring facts about two people's relationship from a brief clip of how they
interact, built on the Seamless Interaction
dataset. Each released set is a config of this repository.
🎧 Multi-modal — text, audio, and video for every clip
🎯 Objective label —… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/friend-bench.serena-synthetic-it-27h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.mcl-mmcl-audiocapsid-en-codeswitch-dataset-alternative
Indonesian–English Code-Switching Synthetic Speech Dataset
Synthetic speech generated for the undergraduate final project "Handling
Code-Switching in Automatic Speech Recognition for Low-Resource Language
Pairs: An Indonesian–English Case Study", School of Electrical Engineering
and Informatics, Institut Teknologi Bandung.
This dataset contains synthetic audio produced from the Indonesian–English
code-switching text corpora released in the companion repository below. It
was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.fsd50k-cc0-curated-v1
FSD50K CC0 Curated v1
A 1,408-clip CC0-only subset of FSD50K (Fonseca et al., 2022), curated for an RNN/LSTM audio generation teaching assignment.
Contents
1,408 WAV files from the FSD50K dev split (<file_id>.wav)
fsd50k_cc0_dev_curated_v1_manifest.csv — per-clip metadata
All files are CC0 / public domain — no attribution required
18 primary labels covering music instruments and nature ambient sounds
Total size: ~1.5 GB, total duration: ~4.73 hours
Original sample rates… See the full description on the dataset page: https://huggingface.co/datasets/HughXuechen/fsd50k-cc0-curated-v1.voiceguard-competition
VoiceGuard — Deepfake Audio Detection Competition
Pelatnas IOAI 2026 | Task 3 of 3
Detect whether a 4-second audio clip is real human speech or AI-generated (TTS/deepfake). Submit probability scores — AUROC is the metric.
Task
Input: .wav audio file (4 seconds, 16 kHz mono)Output: score — probability (0–1) that the audio is fakeMetric: AUROC (Area Under ROC Curve)
Dataset
Split
Real
Fake
Total
Train
2,874
2,874
5,748
Test
627
627
1,254… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/voiceguard-competition.The_Arabic_News_speech_Corpus_Dataset
Arabic News Speech Corpus Dataset
This dataset is an Arabic speech corpus that supports the development of syllable-based Arabic speech recognition using Wav2Vec-2 architecture and a 5-gram language model. It consists of Modern Standard Arabic (MSA) syllables extracted from TV news broadcasts, annotated with diacritics.
Dataset Details
Dataset Description
This corpus contains 15 hours of WAV audio recordings transcribed into diacritized Modern Standard Arabic… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimSalah/The_Arabic_News_speech_Corpus_Dataset.conversational-sarcasm-benchmark
Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only
A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language
YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit
pairs a short target utterance with the preceding context that makes its
figurative reading available, and carries a categorical label plus a free-text rationale.
This repository contains no audio. It ships annotations, transcriptions, and the
source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.test4
test4
This is a merged speech dataset containing 345 audio segments from 2 source datasets.
Dataset Information
Total Segments: 345
Speakers: 7
Languages: en
Emotions: neutral, sad, angry, happy
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion (neutral, happy… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test4.Chinese-Speech-Dataset
🎧 Chinese (Simplified) Speech Dataset
The Chinese (Simplified) speech dataset is a high-quality speech audio dataset developed to support scalable AI and machine learning solutions with diverse and structured audio data. It contains 105 hours of speech data across 700 audio files, delivered in MP3 and WAV formats, with a total size of 229 MB. This well-balanced audio dataset provides reliable voice data, featuring 54% female and 46% male speakers, with age distribution ranging from… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chinese-Speech-Dataset.circor-heart-soundhuman-robot-conversation-korean
Human-Robot Conversation Dataset (Korean) - 660+ Hours
Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.clonecommonhuman-robot-conversation-korean
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.YDX07_Multilingual_Corpus_2026crowd-recital-yi
About
This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project.
Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read.
Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below).
The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.Yadonay-YDX07_Multilingual_Corpus_2026human-robot-conversation-german
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.test5
test5
This is a merged speech dataset containing 1806 audio segments from 8 source datasets.
Dataset Information
Total Segments: 1806
Speakers: 47
Languages: en
Emotions: happy, neutral, sad, angry
Original Datasets: 8
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion (neutral… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test5.commonvoice-mnCommonVoiceAz
