datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-YODAS-ENepstractor-raw
Epstractor: Epstein Archives Dataset
A comprehensive archive of documents, images, audio, and video files from multiple Epstein-related releases, including estate records and Department of Justice materials obtained through FOIA requests.
Dataset Description
This dataset contains 59,420 files totaling 115.23 GB from three major document releases, plus 2 large videos (40GB) available via a separate config:
Epstein Estate 2025-09: 5 files, 0.09 GB
Epstein Estate 2025-11:… See the full description on the dataset page: https://huggingface.co/datasets/public-records-research/epstractor-raw.DeepDialogue-orpheus
DeepDialogue-orpheus
DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text.
🚨 Important Notice
This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.TASTE-DumpEmilia-ENJailbreak-AudioBenchJailbreak-AudioBench-Plusmu-bench
Dataset Card for μ-Bench (Leaderboard | Code)
μ-Bench is a multilingual transcription benchmark built from real customer-service phone conversations.
Dataset Details
Dataset Description
Most public ASR benchmarks are either English-only or built from read speech in quiet studios. μ-Bench fills that gap: real phone-call audio, five languages, and metrics that go beyond Word Error Rate to distinguish meaning-changing errors from surface-level… See the full description on the dataset page: https://huggingface.co/datasets/sierra-research/mu-bench.Bambara-Speech-Translation-Data
AfVoices-Translated (Bambara-English)
This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks.
Methodology
We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository.
Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.beans_watkins
Dataset Card for "beans_watkins"
Dataset Description
Paper:
https://doi.org/10.1121/2.0000358
Dataset Summary
This dataset contains annotated recordings of marine mammal sounds with splits and preprocessing like described in BEANS. It is used for classification tasks.
Data Splits
train
train_low
valid
test
1017
203
339
339
DBD-research-group
Dataset Card for "DBD-research-group"
More Information needed
AST-Music-Data-45KAST-Music-Data-82KDeepDialogue-xtts
DeepDialogue-xtts
DeepDialogue-xtts is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions.
This repository contains the XTTS-v2 variant of the dataset, where speech is generated using XTTS-v2 with explicit emotional conditioning.
🚨 Important
This dataset is large (~180GB) due to the inclusion of high-quality audio files. When cloning the… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-xtts.inat_soundsformosaspeech
Formaosa Speech: Traditional Chinese Long-form Speech
Derived from https://scidm.nchc.org.tw/dataset/grandchallenge .
Emilia-YODAS-DEWolof-ASR-DataA curated Wolof ASR dataset from various sources:
Split
Fleurs
Alfa
CV
Kallama
UB
Total
Train
8.72
16.13
34.97
33.60
4.52
97.94
Test
1.75
2.84
6.21
5.91
1.12
17.83
This dataset was used to finetune Wolof-HuBERT-Base for ASR.
laion-tts-annotated-v1-research
LAION TTS Annotated v1 — research subsets
29,739,936 annotated speech utterances across three subsets — with the audio, the codec tokens
and the complete annotation stack.
The audio in this repository comes from podcasts that are openly available on the internet and
consists of short snippets only. We cannot redistribute the audio itself, so it is made available
here for non-commercial research use by collaboration partners within our TTS research.
The other six subsets of this… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1-research.ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples
validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.beans_bats
Dataset Card for "beans_bats"
Dataset Description
Paper: https://doi.org/10.1038/sdata.2017.143
Dataset Summary
This dataset contains annotated recordings of Egyptian fruit bats with splits and preprocessing like described in BEANS. It is used for classification tasks.
Data Splits
train
train_low
valid
test
6000
1200
2000
2000
dataset-from-restoreMOSS-TTS-ky-kk-bench
MOSS-TTS Kyrgyz/Kazakh Cross-Lingual Benchmark
Paired renderings of the same prompts by two text-to-speech models: MOSS-TTS v1.5
as released, and the same model with a Kyrgyz/Kazakh QLoRA adapter. Each row puts
the two side by side, so the effect of the fine-tune can be judged by ear rather than
from a metric.
The grid is deliberately cross-lingual: every reference voice is used with every
language, so an English speaker reads Kyrgyz, a Kazakh speaker reads Russian, and so on.… See the full description on the dataset page: https://huggingface.co/datasets/KaniTTS-research-team/MOSS-TTS-ky-kk-bench.beans_rfcx
Dataset Card for "beans_rfcx"
Dataset Description
Paper: https://doi.org/10.1016/j.ecoinf.2020.101113
Dataset Summary
This dataset contains continuous soundscape
recordings of 24 species of frogs and birds collected by Rain-
forest Connection (RFCx) with splits and preprocessing like described in BEANS. It is used for detection tasks.
Data Splits
train
train_low
valid
test
2836
964
945
946
AST-Music-Data-1Kemolia_filtered_v1
Emolia Filtered v1 (103,521 samples)
Subset of laion/Emolia processed through
the audio_filter pipeline.
All samples preserved (good + bad + uncertain), with filter results as additional columns.
Pipeline
Stage
Model
Purpose
1. Quality
Dual LogisticRegression (V1 SR<=24kHz / V3 SR>24kHz) on DSP metrics
Detect noise, clipping, robotic, bandwidth-limited audio
2. Speaker
Pyannote ONNX segmentation-3.0
Detect overlapping speakers
Speaker filter runs only on… See the full description on the dataset page: https://huggingface.co/datasets/KaniTTS-research-team/emolia_filtered_v1.Music-Evalbeans_dcase
Dataset Card for "beans_dcase"
Dataset Description
Paper: https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Morfi_52.pdf
Dataset Summary
This is the dataset used for the DCASE 2021 Task and contains annotated mammal and bird multi-species recordings with splits and preprocessing like described in BEANS. It is used for detection tasks.
Data Splits
train
train_low
valid
test
702
151
234
232
Emilia-ZHthaha-research-data2-v2
Nepali Speech Dataset (YouTube-sourced)
441 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 441 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2.
