datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-YODAS-ENreazonspeechlaions_got_talent_rawEmilia-ENbiggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
osu-beatmaps
osu! Beatmaps Dataset (WebDataset)
A collection of ranked/loved osu! beatmaps with audio and chart data, in WebDataset format.
Dataset Variants
Variant
Audio Format
Description
original
MP3/OGG/WAV
Full quality original audio files
compressed
64kbps Mono Opus
Compressed audio for smaller download
from datasets import load_dataset
# Load original audio variant
ds = load_dataset("project-riz/osu-beatmaps", "original", streaming=True)
# Load compressed… See the full description on the dataset page: https://huggingface.co/datasets/project-riz/osu-beatmaps.reazon-speech-v2-clone
Reazon Speech v2 dataset mirror
Original Dataset Source
Hugging Face Dataset Page: reazon-research/reazonspeech
Project Page: Reazon Research
License
This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction:
TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1multi_round_speech_180kRU-AI-noise
RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
This is the noise agumented data for paper: RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
The original dataset is avaliable at zenodo:
https://zenodo.org/records/11406538
The official repo is avaliable at:
https://github.com/ZhihaoZhang97/RU-AI
Reference
We are appreciated the open-source community for the datasets and the models.
Microsoft COCO: Common Objects in… See the full description on the dataset page: https://huggingface.co/datasets/zzha6204/RU-AI-noise.MMAudio-precomputed-results
Precomputed results for MMAudio
Results from four model variants of MMAudio.
All results are in the .flac format with lossless compression.
A cache folder contains the feature caches computed by the evaluation script.
Code: https://github.com/hkchengrex/MMAudio
Evaluation: https://github.com/hkchengrex/av-benchmark
VGGSound
Contains the VGGSound test set results. There are 15220 videos, collected with our best effort. Not all videos in the test sets are available… See the full description on the dataset page: https://huggingface.co/datasets/hkchengrex/MMAudio-precomputed-results.dns5-16k
DNS5 16kHz
Resampled subset of the ICASSP 2022 DNS Challenge dataset.
All audio files resampled from 48kHz to 16kHz and stored as FLAC (lossless compression),
packed into tar shards.
Structure
clean/shard_0000.tar # Clean speech (VCTK and other corpora)
clean/shard_0001.tar
...
noise/shard_0000.tar # Environmental noise (AudioSet, Freesound)
...
impulse_responses/shard_0000.tar # Room impulse responses
...
Each tar contains FLAC files with their… See the full description on the dataset page: https://huggingface.co/datasets/richiejp/dns5-16k.reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.talent_plus_rl_groups_of_50_with_audiobox_scoresnew-rl-vitacommon_voice_21_ru
Dataset Description
Набор данных validated.tsv отфильтрованный по down_votes = 0
📊 Статистика датасета
Информация по сплитам
🔹 Тренировочный набор (train)
Метрика
Значение
Количество семплов
93,531
Общая продолжительность
132.25 часов (476,089.70 секунд)
Средняя продолжительность семпла
5.09 секунд
🔹 Валидационный набор (validate)
Метрика
Значение
Количество семплов
38,836
Общая продолжительность
55.21… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/common_voice_21_ru.Emilia-YODAS-DEru-book-mix-10h
ru-book-mix-10h
A 10-hour synthetic Russian-audiobook diarization benchmark. 600 one-minute
FLAC clips (16 kHz mono, 16-bit, lossless) with NIST RTTM ground truth, generated by
mexus/diarization-benchmark
from its5Q/biggest-ru-book (speech) and bilguun/musan-noise (background).
Intended use: diarization evaluation only. This dataset is not
suitable for training — the same source voices repeat across files, so any
model that trains on it will leak voice identity into its test split.… See the full description on the dataset page: https://huggingface.co/datasets/mexus/ru-book-mix-10h.music-fingerprint-dataset
Neural Audio Fingerprint Dataset
(c) 2021 by Sungkyun Chang
https://github.com/mimbres/neural-audio-fp
This dataset includes all music sources, background noise and impulse-reponses
(IR) samples that have been used in the work ["Neural Audio Fingerprint for
High-specific Audio Retrieval based on Contrastive Learning"]
(https://arxiv.org/abs/2010.11910).
Format:
16-bit PCM Mono WAV, Sampling rate 8000 Hz
Description:
/
fingerprint_dataset_icassp2021/… See the full description on the dataset page: https://huggingface.co/datasets/arch-raven/music-fingerprint-dataset.filimo-farsi-rawlibritts-r-webdatasetOfficial website: https://www.openslr.org/141/
This repository contains LibriTTS-R converted to a WebDataset. The original Wave files have been converted to 64kbps MP3 files for efficient streaming.
LibriTTS-R (paper) is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, published in 2019. The constituent samples of LibriTTS-R are identical to those of LibriTTS, with only the… See the full description on the dataset page: https://huggingface.co/datasets/lucasnewman/libritts-r-webdataset.CompA-R-BackupOriginally from https://huggingface.co/papers/2406.11768, we downloaded it from Google Drive and converted it to HuggingFace, as the train_audio portion was observed to have disappeared.
ruslan-stressed
RUSLAN with Word Stress Marks · RUSLAN с проставленными ударениями
English / Русский
English
What is this?
A drop-in replacement for the metadata of the RUSLAN Russian single-speaker
TTS corpus, with word-stress marks added to every multi-syllabic Russian
word in the transcripts. Audio is bundled unchanged.
The motivation is to train Russian TTS models (e.g. Kokoro, Tacotron, VITS,
StyleTTS, XTTS) that pronounce words with correct lexical stress.
Vanilla… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed.Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.kenya-philippines-twospeaker-english-dialogue
Kenya/Philippines English Dialogue
Two-speaker dialogues in English, recorded on split tracks.
Changelog
Jan 2026: v1 release - vad-segmented WebRTC tracks
Specs
Speakers: >150; ~15 PH, remaining KE
Total duration: ~65 hours
Files sample rate: 48kHz
Actual sample rate: TBD
Language: English (PH, KE accents)
Topics: day-to-day conversation
Collection method
The dataset is built to capture the variety in the Kenyan accent.
The Philippino interviewers… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/kenya-philippines-twospeaker-english-dialogue.AVQA-R1-6KThis repository contains data presented in EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.
For training and inference, please refer to the Code: https://github.com/HarryHsing/EchoInk
Data Format in AVQA-R1-6K:
{
"problem_id": 0,
"problem": "What is the source of the sound in the video?",
"data_type": "image_audio",
"problem_type": "multiple choice",
"options": [
"A. motorcycle",
"B. automobile"… See the full description on the dataset page: https://huggingface.co/datasets/harryhsing/AVQA-R1-6K.rudevices
📊 Сводная статистика аудио-датасетов
📈 Общая статистика по всем датасетам
Метрика
Значение
Всего датасетов/сабсетов
2
Всего семплов
296,394
Общая продолжительность
369.14 часов (1328901.86 секунд)
Средняя продолжительность семпла
4.48 секунд
Распределение объема данных по датасетам
ru_audiobooks_devices ███████████████████████████ 68.5%
rudevices_audio_records ████████████ 31.5%
Датасет: rudevices_audio_records… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/rudevices.rohingya_asr_audioThis is the first public Rohingya language ASR dataset in AI history.
Overview
This dataset contains broadcast audio recordings from the Voice of America (VOA) Rohingya Service. Each file represents a daily news segment, typically 30 minutes in length, automatically segmented into chunks of 5–15 seconds for use in self-supervised ASR, pretraining, language identification, and more.
The content was aired publicly as part of VOA’s Rohingya-language radio program and is therefore… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rohingya_asr_audio.libritts-r-gzmore-synthetic-vocalbursts-raw
More Synthetic Vocal Bursts (Raw)
Synthetic vocal burst audio samples generated from a taxonomy of 202 vocal burst types across multiple text-to-audio and TTS models. Each sample is a short (3–10 second) non-speech vocalization — laughs, cries, gasps, sighs, growls, etc. — generated from text prompts describing the burst type, gender, and age group.
Models Used
Model
Type
Samples
Sample Rate
Notes
DramaBox (ResembleAI/Dramabox)
TTS DiT
2000
44.1 kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/more-synthetic-vocalbursts-raw.
