datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_17_0audiofolder_two_configs_in_metadatamultilingual-speech-commands-15lang
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.cml-tts
Dataset Card for CML-TTS
Dataset Summary
CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG).
CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.audiofolder_single_config_in_metadatacovost2This is a partial copy of CoVoST2 dataset.
The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer.
The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger.
As such, not all the data is included: Only the validation and test subsets are available.
From the XX_EN subsets, only fr, es, and zh-CN are included.
audiofolder_no_configs_in_metadataMultitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus.
MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching.
ASR: Automatic Speech Recognition
SQA: Speech Question Answering
SDS: Spoken Dialogue Summarization
PQA: Paralinguistic Question Answering
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.audiofolder_two_configs_in_metadataSynStard-1000
SynStard-1000
Dataset Summary
SynStard-1000 is a 1,000-hour synthetic dataset for training and evaluating end-to-end speech-to-speech translation (S2ST) models. It is built from English-Chinese parallel texts in the WMT News Commentary v18 corpus and contains approximately 390,000 sentence pairs with paired synthetic speech.
Dataset Structure
.
├── map/
│ └── all.tsv
│── text/
│ ├── en/
│ │ ├── en.txt
│ │ ├── en_1.txt
│ │ ├── ...
│ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/cksqs/SynStard-1000.ciCLAP_freesound
LAION-Audio-630K Freesound Dataset
LAION-Audio-630K is the largest audio-text dataset publicly available and a magnitude larger than previous audio-text datasets (by 2022-11-05). Notably, it combines eight distinct datasets, which includes the Freesound dataset.
Specifically, this Hugging face repository contains two versions of Freesound dataset. Details of each dataset (e.g. how captions are made etc.) could be found in the "datacard" column of the table below.
Freesound (full):… See the full description on the dataset page: https://huggingface.co/datasets/Meranti/CLAP_freesound.Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
satb-choral-dataset
SATB Choral Source Separation Dataset (Compressed)
This dataset contains preprocessed 4-second audio chunks from the Choral Singing Dataset (CSD) formatted for SATB (Soprano, Alto, Tenor, Bass) voice source separation tasks.
This is the compressed version with 8kHz sample rate and int8 precision for smaller file sizes.
Data Structure
Folder Structure
├── chunks/ # All individual chunk .pt files
├── quality_samples/ # Sample WAV… See the full description on the dataset page: https://huggingface.co/datasets/EwanB/satb-choral-dataset.c-sac-corpora
C-SAC LibriTTS-R training subset
Deterministically selected and resampled speech from
mythicinfinity/libritts_r
for the C-SAC causal speech-codec program. The package retains source revision,
Parquet shard, row, utterance, transcript, and content hashes. LibriTTS-R is
distributed under CC BY 4.0; downstream users remain responsible for attribution.
Only prefixes with a hash-bound _COMPLETE.json sentinel are admissible.
audiofolder_two_configs_in_metadata_with_defaultYO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.CASTLE2024
What is CASTLE?
The CASTLE dataset is a large-scale, multimodal dataset designed for advancing research in lifelogging, human activity recognition, and multimodal retrieval. It provides a rich collection of time-aligned sensor and video data for analysis and benchmarking. See the Paper (or its arXiv pre-print) for more details.
You can check our website for more details.
Characteristics
Captured over four days in a controlled environment
10 participants engaged… See the full description on the dataset page: https://huggingface.co/datasets/CASTLE-Dataset/CASTLE2024.AnyAudio-Judge-Corpus
AnyAudio-Judge Corpus
An SFT training corpus that powers the AnyAudio-Judge evaluator. Each sample contains:
An audio clip (referenced relatively under audios/).
A multi-turn chat (messages) where the user enumerates a list of decomposed binary rubric items and the assistant answers them in JSON, with per-item evidence (Chain-of-Thought rationale).
A coarse label ("yes" if the caption originally matched the audio, "no" otherwise) and a tag describing how the caption was… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Corpus.coral-v3
CoRal: Danish Conversational and Read-aloud Dataset
Version 3.0
Dataset Overview
CoRal is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations.
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/CoRal-project/coral-v3.telegram-audiobook-chizzled
Telegram Persian Audiobook Chizzled
1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release
This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.Multitask-National-Speech-Corpus-v1-extendEarnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.uyghur-common-voice-tts
Uyghur Common Voice TTS Dataset
A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice.
Dataset Summary
Property
Value
Language
Uyghur (ug)
Total Samples
43,054
Train Samples
40,901
Validation Samples
2,153
Audio Format
WAV
Source
Mozilla Common Voice
License
CC0-1.0
Dataset Structure
/
├── train.jsonl # Training data (40,901 samples)
├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.UniLSTalkDataset
UniLS-Talk Dataset
To enable research on unified speaking and listening avatar generation, we curate and construct the UniLS-Talk Dataset, a large-scale collection of high-quality 3D facial motion data. We apply a carefully designed tracking pipeline to extract per-frame FLAME parameters, including expression coefficients, eye-gaze, jaw pose and head pose annotations. The dataset comprises two complementary parts:
Paired conversational data sourced from the Seamless Interaction… See the full description on the dataset page: https://huggingface.co/datasets/xg-chu/UniLSTalkDataset.Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
m4singerForDiffSingerSVSvibravox
Dataset Card for VibraVox
👀 While waiting for the TooBigContentError issue to be resolved by the HuggingFace team, you can explore the dataset viewer of vibravox-test
which has exactly the same architecture.
DATASET SUMMARY
The VibraVox dataset is a general purpose audio dataset of french speech captured with body-conduction transducers.
This dataset can be used for various audio machine learning tasks :
Automatic Speech Recognition (ASR) (Speech-to-Text… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/vibravox.cv_corpus_v22
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.zoengjyutgaai
張悦楷講古語音數據集
English
呢個係張悦楷講《三國演義》、《水滸傳》、《走進毛澤東的最後歲月》、《鹿鼎記》語音數據集。張悦楷係廣州最出名嘅講古佬 / 粵語説書藝人。佢從上世紀七十年代開始就喺廣東各個收音電台度講古,佢把聲係好多廣州人嘅共同回憶。本數據集收集嘅係佢最知名嘅四部作品。
數據集用途:
TTS(語音合成)訓練集
ASR(語音識別)訓練集或測試集
各種語言學、文學研究
直接聽嚟欣賞藝術!
TTS 效果演示:https://huggingface.co/spaces/laubonghaudoi/zoengjyutgaai_tts
説明
所有文本都根據 https://jyutping.org/blog/typo/ 同 https://jyutping.org/blog/particles/ 規範用字。
所有文本都使用全角標點,冇半角標點。
所有文本都用漢字轉寫,無阿拉伯數字無英文字母
所有音頻源都存放喺/source,為方便直接用作訓練數據,切分後嘅音頻都放喺 opus/
所有 opus 音頻皆為 48000… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/zoengjyutgaai.
