datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_17_0xperience-10m
⚠️ Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified.
Interactive Intelligence from Human Xperience
Xperience-10M
Dataset Summary
Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial… See the full description on the dataset page: https://huggingface.co/datasets/ropedia-ai/xperience-10m.covost2This is a partial copy of CoVoST2 dataset.
The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer.
The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger.
As such, not all the data is included: Only the validation and test subsets are available.
From the XX_EN subsets, only fr, es, and zh-CN are included.
IndicVoices
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Updates
[23 December 2025] We now have 11,200 hours of transcribed data! 🎉
Overview
INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.AISHELL-3AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It can be used to train multi-speaker Text-to-Speech (TTS) systems.The corpus contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers and total 88035 utterances. Their auxiliary attributes such as gender, age group and native accents are explicitly marked and provided in the corpus. Accordingly, transcripts in… See the full description on the dataset page: https://huggingface.co/datasets/AISHELL/AISHELL-3.AISHELL-4
AISHELL-4
Identifier: SLR111
Summary: A Free Mandarin Multi-channel Meeting Speech Corpus, provided by Beijing Shell Shell Technology Co.,Ltd
Category: Speech
License: CC BY-SA 4.0
Downloads (use a mirror closer to you):
train_L.tar.gz [7.0G] ( Training set of large room, 8-channel microphone array speech
) Mirrors:
[US]
[EU]
[CN]
train_M.tar.gz [25G] ( Training set of medium room, 8-channel microphone array speech
) … See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-4.malaysian-youtube
Malaysian Youtube
Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours.
URLs data at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube/data
Notebooks at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube
How to load the data efficiently?
import pandas as pd
import json
from datasets import Audio
from torch.utils.data import DataLoader, Dataset
chunks = 30
sr = 16000
class Train(Dataset):… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-youtube.AIR-Bench-Dataset
AIR-Bench
Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks.
The former consists of 19 tasks with approximately 19k single-choice questions.
The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon).
Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.Rasa
Rasa: Towards Building an Expressive Multilingual Text-To-Speech Dataset for Indian Languages
Funded by: Bhashini, Ministry of Electronics and Information Technology, Government of IndiaSupported by: EkStep Foundation and Nilekani Philanthropies
Overview
We introduce Rasa, the first high-quality multilingual expressive Text-to-Speech (TTS) dataset for any Indian language. It comprises a minimum of 20 hours per speaker with a target of covering
a female and male… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rasa.audio_samples_1kindicvoices_r
IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS
Dataset Summary
IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.AISHELL-4genshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.VoiceAgentBench
VoiceAgentBench
This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978).
VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.Kathbath
Kathbath
Kathbath is an human-labeled ASR dataset containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India
Languages
Bengali
Gujarati
Kannada
Hindi
Malayalam
Marathi
Odia
Punjabi
Sanskrit
Tamil
Telugu
Urdu
Licensing Information
The IndicSUPERB dataset is released under this licensing scheme:
We do not own any of the raw text used in creating this dataset.
The text data… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Kathbath.AISHELL-3
AISHELL-3
Identifier: SLR93
Summary: Mandarin data, provided by Beijing Shell Shell Technology Co., Ltd.
Category: Speech
License: Apache License v.2.0
Downloads (use a mirror closer to you):
data_aishell3.tgz [19G] (speech data and transcripts
) Mirrors:
[US]
[EU]
[CN]
About this resource:AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus
published by Beijing Shell Shell Technology Co.,Ltd. It can be… See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-3.everyayah﷽
Dataset Card for Tarteel AI's EveryAyah Dataset
Dataset Summary
This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in Arabic.
Dataset Structure
Data Instances
A typical data point comprises the audio file audio, and its transcription called text.
The duration… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/everyayah.ZeroSpeech
ZeroSpeech
A large synthetic Vietnamese speech corpus for ASR training: 9,867,987
utterances / 26,896 hours, spoken by 199,265
distinct voices, generated with
ZeroTTS from web and
conversational text.
Every clip is 16 kHz mono FLAC, 1–30 s, paired with the exact text it was
synthesized from.
Fields
field
type
description
audio
Audio(16 kHz)
the waveform, FLAC-encoded
text
string
the transcript — the exact string given to the TTS
source
string
which… See the full description on the dataset page: https://huggingface.co/datasets/zeroweight-ai/ZeroSpeech.turn-benchmark-dev
TurnBench - Dev Set
TurnBench is a benchmark for evaluating
conversational turn-taking: end-of-turn and interruption detection on real
annotated two-speaker conversations.
This repository contains the development split: 38 English conversations,
about 7.3 hours of audio, packaged as one row per conversation. Each row contains
two time-aligned per-speaker audio streams plus three independent annotator
tracks per speaker.… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-dev.EA-DIAIMEfrom datasets import load_dataset
dataset = load_dataset('disco-eth/AIME')
AIME: AI Music Evaluation Dataset
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset.
The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset.
The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.music-ai-human-test-audio
Interpretable AI and Human Music Evaluation Archive
Research audio and versioned experiment outputs for an English graduation thesis.
The audio archive is incomplete. Completed experiments and verified partial
audio publications must not be confused with whole-project delivery completion.
No blanket license is assigned to this mixed-source archive.
Completed experiments and thesis
The BC extension, expanded YuE Native30 evaluation, locked YuE Native30 scoring… See the full description on the dataset page: https://huggingface.co/datasets/EZMONYI/music-ai-human-test-audio.suno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.Shrutilipi
Shrutilipi
Overview
Shrutilipi is a labelled ASR corpus obtained by mining parallel audio and text pairs at the document scale from All India Radio news bulletins for 12 Indian languages: Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Sanskrit, Tamil, Telugu, Urdu. The corpus has over 6400 hours of data across all languages.
This work is funded by Bhashini, MeitY and Nilekani Philanthropies
Usage
The datasets library… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Shrutilipi.gigaspeechsmart-turn-data-v3.2-trainTraining dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train.himalaya-ai-stt-datasetTimeGround-1M
TimeGround-1M
Synthetic English audio dataset for time-aware speech understanding, covering temporal localization, temporal description, and timed summaries.
Data Filtering
We use 14k hours of audio from YODAS2 English shards, selected from a 24k-hour source pool after language- and silence-ratio filtering. Synthetic annotations were generated for three time-grounded tasks, then filtered through LLM-based verification, deterministic validity checks, and… See the full description on the dataset page: https://huggingface.co/datasets/ai-sage/TimeGround-1M.captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
ReactNet
ResponseNet
ResponseNet is a large-scale dyadic video dataset designed for Online Multimodal Conversational Response Generation (OMCRG). It fills the gap left by existing datasets by providing high-resolution, split-screen recordings of both speaker and listener, separate audio channels, and word‑level textual annotations for both participants.
Paper
If you use this dataset, please cite:
ResponseNet: A High‑Resolution Dyadic Video Dataset for Online Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/awakening-ai/ReactNet.
