datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
unsupervised_peoples_speech
Dataset Card for Unsupervised Peoples Speech
Dataset Description
Dataset Summary
The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers.
Point of Contact: MLCommons Datasets Discord
Dataset Structure
This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.psg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.UniLSTalkDataset
UniLS-Talk Dataset
To enable research on unified speaking and listening avatar generation, we curate and construct the UniLS-Talk Dataset, a large-scale collection of high-quality 3D facial motion data. We apply a carefully designed tracking pipeline to extract per-frame FLAME parameters, including expression coefficients, eye-gaze, jaw pose and head pose annotations. The dataset comprises two complementary parts:
Paired conversational data sourced from the Seamless Interaction… See the full description on the dataset page: https://huggingface.co/datasets/xg-chu/UniLSTalkDataset.Real_Voicequranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.FOR-normdanish-asr-unified
Danish ASR Unified Dataset
Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours):
Source
Samples
Description
VoxPopuli
1,775,578
European Parliament recordings
ftspeech
995,677
Danish Parliament (Folketinget)
CoRal-v3 read_aloud
299,255
Read-aloud Danish speech
nst-da
182,605
NST Danish speech
CoRal-v3 conversation
147,249
Conversational Danish speech
nota
98,600
Danish broadcast media
Common Voice 17
3,484
Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.ahead_ds_unmixed
Another HEaring AiD DataSet (AHEAD-DS) unmixed
Another HEaring AiD DataSet (AHEAD-DS) unmixed is an audio dataset labelled with audiologically relevant scene categories for hearing aids. This dataset contains the environment and speech sounds before they were mixed. The file ahead_ds_unmixed.csv documents the details of every file.
Website
Paper
Code
Dataset AHEAD-DS
Dataset AHEAD-DS unmixed
Models
Description of data
All files are encoded as single channel WAV, 16 bit… See the full description on the dataset page: https://huggingface.co/datasets/hzhongresearch/ahead_ds_unmixed.majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.farsi-asr-unified-cleaned
🎧 Farsi ASR Unified Dataset (Parquet Sharded Edition)
Overview
The Farsi ASR Unified Dataset is a large-scale, high-quality, and fully standardized collection of Persian (Farsi) speech-to-text data — designed specifically for modern machine learning and ASR (Automatic Speech Recognition) workflows.
This dataset consolidates audio–text pairs from multiple open sources, applies a rigorous cleaning and normalization pipeline, and stores everything efficiently in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/kiarashQ/farsi-asr-unified-cleaned.UNO-Bench UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
🔔News
🔥[2025/12/04] We have released the evaluation scripts uno-eval, a unified evaluation framework for omni-modal benchmarks. More benchmarks will be supported in the future.
🔥[2025/12/04] We have released the scoring model UNO-Scorer-Qwen3-14B. Feel free to use it!
👀 UNO-Bench Overview
Multimodal Large Languages models have… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/UNO-Bench.maestro-unidac4-ytsv
MAESTRO + ASAP audio and MIDI tokens (U-MusT)
Tokenized MAESTRO v3.0.0 for
U-MusT: DAC audio tokens and MT3-style MIDI event arrays,
covering roughly 199 hours of Disklavier-captured piano performance with precisely aligned MIDI.
This repository also contains ASAP-derived data. lmx/ and asap_note_events/ come from the
ASAP dataset, whose audio is itself MAESTRO. Both
carry the same license, so nothing conflicts, but the repository name mentions only one of the two
corpora it… See the full description on the dataset page: https://huggingface.co/datasets/malerlab/maestro-unidac4-ytsv.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1unified-kannada-asr-1.0
Dataset Card for "unified-kannada-asr-1.0"
More Information needed
Whisper-fine-tune-2hailuo-ai-voices
Hailuo AI Voices Dataset 🎤
A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks.
📊 Dataset Overview
The dataset provides a comprehensive collection of voice samples with the following features:
Feature
Description
Audio Files
High-quality WAV format recordings
Transcription
Accurate transcriptions of each… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-voices.MUSAN-speech_unit_part2
Dataset Card for "MUSAN-speech_unit_part2"
More Information needed
bpsd-unirqvae3-unidac4-ytsv
BPSD score-image, audio and notation tokens (U-MusT)
Tokenized Beethoven Piano Sonata Dataset v2 for
U-MusT — the test-only split, and the only corpus in the
collection carrying all four modalities: score-image tokens, audio tokens, and LMX notation.
Because it is held out for evaluation, the image tokens here are not shift-augmented: they have
shape (1, 1, H, W, 4), a single tokenization. The audio tokens retain the 9-variant stack.
BPSD ships no system-level image alignment… See the full description on the dataset page: https://huggingface.co/datasets/malerlab/bpsd-unirqvae3-unidac4-ytsv.iv_speaker_disjoint_sociodem_unaware_dsMUSAN-speech_unit_part1
Dataset Card for "MUSAN-speech_unit_part1"
More Information needed
MUSAN-noise_unit_part2
Dataset Card for "MUSAN-noise_unit_part2"
More Information needed
UnityShotsBench
UnityShots Benchmark
A multilingual, multi-cultural k-shot storytelling benchmark for evaluating multi-shot
audio-video generation. Each case is a short cinematic story told across several shots, with a
consistent cast whose identity, voice, and world must persist across every cut.
This is the evaluation benchmark released with UnityShots: Memory-Driven Multi-Shot
Audio-Video Generation with Boundary-Aware Gating.
📄 Paper: arXiv:2606.21661
🌐 Project page:… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/UnityShotsBench.MUSAN-music_unit_part1
Dataset Card for "MUSAN-music_unit_part1"
More Information needed
ASVSpoof21_PA2UnityShotsBench
UnityShots Benchmark
A multilingual, multi-cultural k-shot storytelling benchmark for evaluating multi-shot
audio-video generation. Each case is a short cinematic story told across several shots, with a
consistent cast whose identity, voice, and world must persist across every cut.
This is the evaluation benchmark released with UnityShots: Memory-Driven Multi-Shot
Audio-Video Generation with Boundary-Aware Gating.
📄 Paper: arXiv:2606.21661
🌐 Project page:… See the full description on the dataset page: https://huggingface.co/datasets/zxc0135/UnityShotsBench.United-Syn-Meduniversebench
UniVerseBench
The evaluation split of UniVerse (同谣).Training data lives in UniVerseSet.
UniVerseBench is a multilingual folk-music understanding benchmark for large audio–language models (LALMs). It asks models to listen, not to guess from language priors.
「诗言志,歌永言,声依永,律和声。」—《尚书·舜典》
Sister dataset (training)
universe-team/universeset
Live museum demo
http://143.89.224.8:8790/
What's here
Two views of the same benchmark:
Subset… See the full description on the dataset page: https://huggingface.co/datasets/universe-team/universebench.hailuo-ai-jokes
Hailuo AI Jokes Dataset 🎤
A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks.
🎙️ Dataset Content
The dataset contains a diverse set of synthetic voice recordings generated by Hailuo AI Audio. The texts are sourced from a variety of public domain jokes and humorous anecdotes. Each audio sample is accompanied by the… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-jokes.MUSAN-music_unitlibrispeech_unit_speech
Dataset Card for "librispeech_unit_speech"
More Information needed
