datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BEAT2EK100
Motivation
The actual download link is very slow, including the academic torrent. Therefore, to spare fellow community members from this misery, I am uploading the dataset here.
Source
You can fnd the original source to download the dataset: https://github.com/epic-kitchens/epic-kitchens-download-scripts
Citation
@INPROCEEDINGS{Damen2018EPICKITCHENS,
title={Scaling Egocentric Vision: The EPIC-KITCHENS Dataset},
author={Damen, Dima and Doughty, Hazel and… See the full description on the dataset page: https://huggingface.co/datasets/itruonghai/EK100.MADB-Dataset
MADB: Music Aesthetics Dataset and Benchmark
Dataset Description
MADB is a large-scale dataset for music aesthetic evaluation, designed to support research on multi-dimensional and subjective music perception.
The dataset contains approximately 10,000 music tracks, each annotated by multiple trained annotators across 10 perceptual dimensions and one overall score. In addition, each track includes textual comments and semantic tags (genre and mood), enabling… See the full description on the dataset page: https://huggingface.co/datasets/sirui1/MADB-Dataset.maleo-short-1.5H
Dataset Card for Maleo Short 1.5H
Dataset Description
Dataset Summary
Maleo Short 1.5H is a manually curated, rigorously annotated speaker diarization dataset designed to benchmark State-of-the-Art (SOTA) models against complex, "in-the-wild" media domains. While modern diarization pipelines excel in controlled acoustic environments (like telephony or reading corpora), they heavily struggle with the overlapping speech, sound effects, and rapid speaker shifts… See the full description on the dataset page: https://huggingface.co/datasets/maleo-ai/maleo-short-1.5H.Hadou-Voice-Dataset
Hadou Voice Dataset
ハドウ本人が収録した、日本語音声データセットです。
このページで、特徴の異なる2種類のデータセットを公開しています。
配布データ
設定名
内容
音声数
合計時間
v1(おすすめ)
Hadou Calm Voice Dataset v1。落ち着いた中音域、AIキャラクター向けボイスが多めの音声データ
966
約114.02分
v0
Hadou ITA Corpus Dataset v1。ITAコーパスを読み上げた自然な話し声
424
約38.95分
v1 には、AICAコーパス500文、ITAコーパス324文、感情・態度付き90文、同文異演技40文、強度段階12文を収録しています。
v1の詳細: v1/README.txt
v0の詳細: v0/README.txt
読み込み例
from datasets import load_dataset
# 新しい966音声(既定)
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/hadou1225/Hadou-Voice-Dataset.TUT2018-ov1africanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
Phone_Timings_Database
📖 TajweedAI: Quranic Phoneme Timing Benchmark (Phases 1, 2 & 3)
📌 Project Overview
TajweedAI evaluates Quranic recitation accuracy by analyzing both pronunciation (phoneme classification) and timing (rule duration evaluation).
This benchmark provides empirical, tempo-normalized duration boundaries for all 70 Quranic phonemes derived from forced alignments (MFA trained on Quranic audio) across 7 master reference reciters:
Sheikh Mahmoud Khalil Al-Husary (Gold… See the full description on the dataset page: https://huggingface.co/datasets/AhmedTamertechno1/Phone_Timings_Database.mcl-mmcl-audiocapsfsd50k-cc0-curated-v1
FSD50K CC0 Curated v1
A 1,408-clip CC0-only subset of FSD50K (Fonseca et al., 2022), curated for an RNN/LSTM audio generation teaching assignment.
Contents
1,408 WAV files from the FSD50K dev split (<file_id>.wav)
fsd50k_cc0_dev_curated_v1_manifest.csv — per-clip metadata
All files are CC0 / public domain — no attribution required
18 primary labels covering music instruments and nature ambient sounds
Total size: ~1.5 GB, total duration: ~4.73 hours
Original sample rates… See the full description on the dataset page: https://huggingface.co/datasets/HughXuechen/fsd50k-cc0-curated-v1.vox1-veri-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
133777
14865
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
vox1-iden-3s
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Identification Split
train
validation
test
# of speakers
1251
1251
1251
# of samples
306208
14479
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
clonevox1-iden-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Identification Split
train
validation
test
# of speakers
1251
1251
1251
# of samples
138361
6904
8251
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
stanislau-shumski-u-bitvakh-i-viaznitsakh-uspaminy-pra-1812-1848-gady-siargei-du
У бітвах і вязьніцах. Успаміны пра 1812–1848 гады
Metadata
Author: Станіслаў Шумскі
Title: У бітвах і вязьніцах. Успаміны пра 1812–1848 гады
Narrator: Сяргей Дубавец
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/stanislau-shumski-u-bitvakh-i-viaznitsakh-uspaminy-pra-1812-1848-gady-siargei-du.EgoAVU_data
[CVPR2026] EgoAVU
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding
See our github for the code and setup instructions.
Check out our homepage and paper for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual understanding. EgoAVU enriches existing egocentric narrations by integrating human actions with environmental context, explicitly linking visible objects and the sounds produced during interactions… See the full description on the dataset page: https://huggingface.co/datasets/jun111111/EgoAVU_data.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.dwesui-grupa-1-neurologia
NeuroSpeechPL
Publiczny eksport HuggingFace zawiera wyłącznie redystrybuowalne audio source=natural. Wiersze TTS są celowo wyłączone z publicznego zbioru danych, ponieważ ich source_license zabrania redystrybucji audio. Pełna lokalna ewaluacja opisana w raporcie korzystała zarówno z nagrań naturalnych, jak i TTS.
Repozytorium zbioru danych HF: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Repozytorium kodu:… See the full description on the dataset page: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia.nps-liminal-soundscapes-v0-1
nps-liminal-soundscapes-v0-1
Private Sonic-Forage Stable Audio soundscape training dataset built from verified National Park Service public-domain sound pages.
License receipt
Primary source pages:
https://www.nps.gov/subjects/sound/gallery.htm — page states files are public domain and may be downloaded; credit requested.
https://www.nps.gov/yell/learn/photosmultimedia/soundlibrary.htm — page states files were recorded in the park, are public domain, and may be used… See the full description on the dataset page: https://huggingface.co/datasets/Sonic-Forage/nps-liminal-soundscapes-v0-1.swahili_10hdarija-tts
Moroccan Darija TTS Dataset
This dataset contains Moroccan Darija speech recordings and their corresponding transcriptions, intended for fine-tuning text-to-speech models.
Dataset Structure
wavs_16k/: Directory containing 16kHz mono 16-bit WAV audio files.
metadata_train.csv: CSV file with training data.
metadata_val.csv: CSV file with validation data.
Usage
Load the dataset using:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Fahd1199/darija-tts.Urdu-Multimodal-Emotion-Datasetdzhordzh-oruel-1984-kupalautsy
1984
Metadata
Author: Джордж Оруэл
Title: 1984
Narrator: купалаўцы
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250 MB.
Each split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/dzhordzh-oruel-1984-kupalautsy.vox1-veri-3s
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
299246
33672
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
tr123
tr123
This is a merged speech dataset containing 11930 audio segments from 24 source datasets.
Dataset Information
Total Segments: 11930
Speakers: 69
Languages: tr
Emotions: sad, neutral, happy, angry
Original Datasets: 24
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr123.Devnagari-ASRtesttr12
testtr12
This is a merged speech dataset containing 165 audio segments from 2 source datasets.
Dataset Information
Total Segments: 165
Speakers: 8
Languages: en
Emotions: happy, angry, neutral
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/testtr12.ds573-trav4cast probe
frisiancommon_voice_19_uk_cropped
