datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.coral-v3
CoRal: Danish Conversational and Read-aloud Dataset
Version 3.0
Dataset Overview
CoRal is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations.
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/CoRal-project/coral-v3.Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
vibravox
Dataset Card for VibraVox
👀 While waiting for the TooBigContentError issue to be resolved by the HuggingFace team, you can explore the dataset viewer of vibravox-test
which has exactly the same architecture.
DATASET SUMMARY
The VibraVox dataset is a general purpose audio dataset of french speech captured with body-conduction transducers.
This dataset can be used for various audio machine learning tasks :
Automatic Speech Recognition (ASR) (Speech-to-Text… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/vibravox.cv_corpus_v22
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.zoengjyutgaai
張悦楷講古語音數據集
English
呢個係張悦楷講《三國演義》、《水滸傳》、《走進毛澤東的最後歲月》、《鹿鼎記》語音數據集。張悦楷係廣州最出名嘅講古佬 / 粵語説書藝人。佢從上世紀七十年代開始就喺廣東各個收音電台度講古,佢把聲係好多廣州人嘅共同回憶。本數據集收集嘅係佢最知名嘅四部作品。
數據集用途:
TTS(語音合成)訓練集
ASR(語音識別)訓練集或測試集
各種語言學、文學研究
直接聽嚟欣賞藝術!
TTS 效果演示:https://huggingface.co/spaces/laubonghaudoi/zoengjyutgaai_tts
説明
所有文本都根據 https://jyutping.org/blog/typo/ 同 https://jyutping.org/blog/particles/ 規範用字。
所有文本都使用全角標點,冇半角標點。
所有文本都用漢字轉寫,無阿拉伯數字無英文字母
所有音頻源都存放喺/source,為方便直接用作訓練數據,切分後嘅音頻都放喺 opus/
所有 opus 音頻皆為 48000… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/zoengjyutgaai.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.central-kurdish-pseudolabel
Central Kurdish → English Pseudo-Labeled Speech Translation Corpus
Dataset Summary
This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish).
The dataset was automatically generated using a pipeline composed of:
Speech segmentation
Automatic Speech Recognition (ASR)
Machine Translation (MT)
The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.WorldSpeech
WorldSpeech
A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Centi234/WorldSpeech.common-voice-13-faThe Persian portion of the original CommonVoice 13 dataset at https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0
Load
# Using HF Datasets
from datasets import load_dataset
dataset = load_dataset("hezarai/common-voice-13-fa", split="train")
# Using Hezar
from hezar.data import Dataset
dataset = Dataset.load("hezarai/common-voice-13-fa", split="train")
majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.apptek_callcenter_dialogues
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents
across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions.
128.6 hours of speech
14 English accent groups
16 service domains
5–15 minute conversations (long-form)
Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.dsb_audio_corpus
Acknowledgements
Thanks to all speakers that contributed to this dataset!
Thanks to "Ludowe Nakładnistwo Domowina" and "Rěčny Centrum WITAJ" for donation of their recordings!
cantonese-radio
Cantonese Radio Pseudo-Transcription Dataset
Contains 14k hours of audio sourced from Archive.org
Columns
order_index: Represents the order of the audio compared to those from the same filename
link: Link of the original full audio
transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding
transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall
used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-radio.cv-v1.0-segment
CommonVoice v1 Phone-Segment Alignments
Phone-level time alignments for 10 languages of Mozilla Common Voice,
packaged in a canonical segmentation schema with embedded 16 kHz audio. The
phone boundaries come from the charsiu/cv_ali
release of MFA alignments; the audio and transcripts come from
Common Voice Corpus 13.0 (2023-03-09).
Dataset summary
lang
train rows
train hrs
val rows
val hrs
test rows
test hrs
en
1,008,669
1,354.0
3,537
4.9
1,285
1.7
rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.gooshkon-chunks
Gooshkon MP3 chunks
A single-column Hugging Face Audio dataset. Every row contains embedded,
playable MP3 bytes in the audio column. Chunks target approximately 30
seconds and are selected at detected quiet intervals with a 12-48 second
safety range.
The source recordings are not transcript-aligned; this release uses acoustic
silence boundaries to avoid cutting through words whenever a usable pause is
available.
Dataset runtime
The dataset contains approximately 2… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/gooshkon-chunks.arabic-audio-collection-algerian-loubna-stories
Loubna Stories Arabic Speech Dataset
Dataset Summary
The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.UniST
UniST
This dataset contains UniST codec-token training data exported from local metadata and codec results.
We train UniSS with UniST data.
Schema
id: sample identifier
transcription: source transcription from metadata text
translation: qwen_trans, falling back to trans_text
source_glm, target_glm: GLM token lists
source_bicodec, target_bicodec: bicodec semantic token lists
bicodec_global: source bicodec global token list
dataset_name, src_lang, tgt_lang, split:… See the full description on the dataset page: https://huggingface.co/datasets/cmots/UniST.ASCEND
Dataset Card for ASCEND
Dataset Summary
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/CAiRE/ASCEND.dinner-party-corpusThis repository contains a reorganized, utterance-focused version of the Dinner Party Corpus, released by Amazon, the Center for Language and Speech Processing (CLSP) and Johns Hopkins University in September 2019.
Description
The following description is provided in arXiv 1909.13447:
We present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment. The corpus was created by recording multiple groups of four Amazon employee… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/dinner-party-corpus.ContextASR-Bench
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.khm-asr-cultural
Khmer ASR Cultural Dataset
134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.YO-CPT-kk
YO-CPT-kk
YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily
quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker,
TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a
punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and
cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the
voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.DAVE-Corpus
DAVE-Corpus (Open Subset)
Overview
DAVE-Corpus (open subset) is a training dataset for blind source separation (BSS)
and two-speaker speech separation on Chinese meeting speech. It is the redistributable
portion of the training pool of DAVE
(arXiv:2608.09288), our system for the ISCSLP 2026
Real-World AVSE Challenge, and is generated end-to-end by the released synthesis
pipeline from three permissively licensed corpora — AliMeeting, AISHELL-4
(speech) and MUSAN… See the full description on the dataset page: https://huggingface.co/datasets/TaurenMountain/DAVE-Corpus.VoxPopuli-Cleaned-AA
VoxPopuli-Cleaned-AA
Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article
VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models.
This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.
