datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.open-bible-speech-african
Open Bible Resources — African Languages
Spoken-audio Bible recordings aligned to verse-level text for 19 African languages —
roughly 1,741 hours of audio across ~552,907 audio–text pairs (~357 GB).
This dataset is the African-language subset of
davidguzmanr/open-bible-resources,
re-hosted here by AfriSpeech to make the African
languages easy to find and use on their own. The audio and text are unchanged from the
source; only the non-African configurations have been removed. All… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/open-bible-speech-african.open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.ace-opencpop-segments
Citation Information
@misc{shi2024singingvoicedatascalingup,
title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing},
author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe},
year={2024},
eprint={2401.17619},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2401.17619},
}
MonsoonASR-Open-ASR-leaderboard-en-IN
Voice Arena Monsoon en-IN (public test)
Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed.
A conversational Indian English ASR test set that records who was speaking, not only what
was said. Every clip carries twelve speaker attributes — gender, age, native district
and state, education, occupation, income band, handset — so a difference between two
systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.parliament
Parliament
Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the
zh_tw split of disco-eth/WorldSpeech.
It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD
parliamentary proceedings.
This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style
text filtering pipeline. Audio is preserved from the upstream dataset and cast
as a Hugging Face Audio(sampling_rate=24000) feature.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/parliament.open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
openslr
madasr
shrutilipi
flerus
kathbath
indictts
ucla
gali
youtube_transcriptions
Dataset Description
A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings.
Use Cases
Automatic Speech Recognition (ASR) for Uzbek
Text-to-Speech (TTS) synthesis for Uzbek
Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS)
Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.MonsoonASR-Open-ASR-leaderboard-hi-IN
Voice Arena Monsoon hi (public test)
Part of the Open ASR Leaderboard, on the Multilingual tab, where a model is ranked only if it supports every selected language.
A conversational Hindi ASR test set that records who was speaking, not only what
was said. Every clip carries twelve speaker attributes — gender, age, native district
and state, education, occupation, income band, handset — so a difference between two
systems can be traced to a group of speakers instead of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-hi-IN.common_voice_25_zh-TW
Common Voice Scripted Speech 25.0 - Chinese (Taiwan)
This repository mirrors the Chinese (Taiwan) (zh-TW) portion of Mozilla Common Voice Scripted Speech 25.0 in Hugging Face datasets format. The audio has been embedded into Parquet and exposed as a Hugging Face Audio feature at 48 kHz.
Official source page: Mozilla Data Collective - Common Voice Scripted Speech 25.0 - Chinese (Taiwan)
Dataset Details
Field
Value
Dataset ID
cmn2g7eaj01fio10769r1m96n… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/common_voice_25_zh-TW.burmese-speech-refined-openslr-80
Burmese Speech Refined OpenSLR-80
Summary
This dataset is a speech dataset developed based on the original OpenSLR Dataset (SLR80), with the text and audio data carefully reviewed and further refined for Burmese language applications.
In the original OpenSLR Dataset, the Burmese text was transcribed based on how the words were pronounced in the corresponding audio recordings. In this dataset, the original audio and text data were used as a reference, and the text… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/burmese-speech-refined-openslr-80.openstt_balalaika
OpenSTT Annotated by Balalaika
[!IMPORTANT]
Official dataset for our INTERSPEECH 2026 paper
"A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563).
Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika.
If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.
Overview
OpenSTT… See the full description on the dataset page: https://huggingface.co/datasets/lab260/openstt_balalaika.stt-sampler-v1
stt-sampler-v1
Licensing: clips inherit their source dataset's license — CC-BY-4.0
for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE
clips (source_dataset column identifies each clip's origin).
A small, balanced, representative multilingual ASR eval sampler for the
OVOS Plugin Arena:
100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32,
one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")).
Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.myanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (OpenSLR-80)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of OpenSLR HuggingFace page.
Original Source
OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.openslr65-tamil
OpenSLR-65 – Tamil Transcribed Speech
Source: https://www.openslr.org/65/
This dataset contains transcribed high-quality audio of Tamil sentences recorded
by volunteers. It is part of the OpenSLR collection of
free speech resources for low-resource languages.
The data was collected via the
Appen (formerly Figure Eight / CrowdFlower) crowdsourcing
platform and is intended for use in training automatic speech recognition (ASR)
and text-to-speech (TTS) systems.
Data… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr65-tamil.openslr-147-hq-Nahuatl
Veracruz Orizaba Nahuatl Endangered Language
Identifier: SLR147
Summary: Audio corpus of Orizaba (Veracruz) Nahuatl speech (Glottocode: oriz1235; ISO 639-3: nlv)
Category: Speech
License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0)
About this resource:
The substantive material of this deposit was gathered over a 13-month period from February 2022 to March 2023.
It comprised 657 files totaling approximately 119 hours, 26 minutes, 59 seconds of material. All but 81… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-147-hq-Nahuatl.openslr-32-hq-SA-languages-Afrikaans
High quality TTS data for four South African languages - Afrikaans
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Afrikaans
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.fleurs_openslr42_mpwtNOTE: If your colab crashes, please use pip install --upgrade --quiet datasets[audio]==3.6.0 to install datasets[audio] version 3.6.0.
This dataset combined google/fleurs, openslr/openslr42, and cleaned seanghay/khmer_mpwt_speech.
Severals processes are executed:
clean up seanghay/khmer_mpwt_speech: manually correct wrong transcriptions over 2058 rows
normalize transcription: remove invisible white space; process ៗ, numbers, currencies, date into khmer text; and separate each word by space… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/fleurs_openslr42_mpwt.DuplexOmni-Data
DuplexOmni Data
This dataset accompanies DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction.
GitHub: MuyeHuang/DuplexOmni
arXiv: 2606.09186
Files
The root directory includes the Writer-Director source files:
inbound.director.jsonl
outbound.director.jsonl
The synthesized parquet training shards are about 9 TB in total, so uploading them is expected to be very slow. They will be uploaded incrementally; each completed… See the full description on the dataset page: https://huggingface.co/datasets/OpenT2S/DuplexOmni-Data.gujarati-f-openslr
Gujarati OpenSLR Female
Interspeech data downloaded from https://www.openslr.org/resources/78/gu_in_female.zip
Dataset Details
Gujarati Data (Most of the entries are <30 seconds and hence Whisper Models can be used for accurate timestamp prediction)
Also, the audio seems to have been spoken by a single female.
openslr80-burmese
OpenSLR-80 – Brumese Transcribed Speech
Source: https://www.openslr.org/80/
This dataset contains transcribed high-quality audio of Burmese sentences recorded
by female volunteers. It is part of the OpenSLR collection of
free speech resources for low-resource languages.
The data was collected via the
Appen (formerly Figure Eight / CrowdFlower) crowdsourcing
platform and is intended for use in training automatic speech recognition (ASR)
and text-to-speech (TTS) systems.… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr80-burmese.openslr42-khmer-male
OpenSLR SLR42 Khmer Male Speech
This dataset is a processed version of the OpenSLR SLR42 Khmer speech dataset.
Dataset Description
This dataset contains approximately 2,906 Khmer speech recordings with corresponding Khmer transcriptions.
Each example contains:
audio: Khmer speech recording
text: Khmer transcription
Dataset Structure
Column
Type
Description
audio
Audio
Khmer speech recording
text
String
Khmer transcription… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/openslr42-khmer-male.openslr-32-hq-SA-languages
SLR32 – High Quality TTS Data for Four South African Languages
Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/
This dataset contains multi-speaker high quality transcribed audio data for four
languages of South Africa: Afrikaans (af_za), Sesotho (st_za),
Setswana (tn_za) and isiXhosa (xh_za).
The dataset consists of WAV files and a TSV file transcribing the audio.
In each folder the file line_index.tsv contains a FileID (which in turn
encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.openslr61-es-ar-full
openslr61-es-ar-full
OpenSLR 61 (Crowdsourced high-quality Argentinian Spanish) consolidado COMPLETO con linaje.
Incluye male + female + weather messages argentinos.
Linaje (trazabilidad por sample)
source: siempre "openslr61"
subset: "main" (frases generales) o "weather" (mensajes de clima)
gender: "m" / "f"
speaker_id: ID anonimizado original del speaker
file_id: ID original del archivo OpenSLR
license: CC-BY-SA-4.0
Schema
campo
tipo… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/openslr61-es-ar-full.openslr42-khmer-tts
OpenSLR42 – High Quality TTS Data for Khmer
This dataset contains high-quality transcribed audio data for Khmer (km-KH).It is the HuggingFace mirror of OpenSLR Resource #42.
Identifier: SLR42
Summary: Multi-speaker TTS data for Khmer
Category: Speech
License: CC BY-SA 4.0
Original source: https://www.openslr.org/42/
Collected by: Google
Copyright: 2016, 2017, 2018 Google LLC
What's inside
Field
Description
filename
Original filename (without… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr42-khmer-tts.openslr-32-hq-SA-languages-Setswana
High quality TTS data for four South African languages - Setswana
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Setswana
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.OpenSLR54-Nepali-ASR-parquet
OpenSLR 54: Large Nepali ASR training data set (unmodified parquet repackaging)
This is an unofficial repackaging of the official OpenSLR 54 release
(SLR54, https://www.openslr.org/54/), converted to parquet so it can be streamed with 🤗 datasets.
It is not affiliated with or endorsed by OpenSLR or the original authors.
All credit for the data belongs to the original creators (see Citation).
What's inside
157,905 utterances, 16 shards: one per original zip… See the full description on the dataset page: https://huggingface.co/datasets/JeevanDai/OpenSLR54-Nepali-ASR-parquet.openslr63
SLR63: Crowdsourced high-quality Malayalam multi-speaker speech data set
This data set contains transcribed high-quality audio of Malayalam sentences recorded by volunteers. The data set consists of wave files, and a TSV file (line_index.tsv). The file line_index.tsv contains a anonymized FileID and the transcription of audio in the file.
The data set has been manually quality checked, but there might still be errors.
Please report any issues in the following issue tracker on… See the full description on the dataset page: https://huggingface.co/datasets/vrclc/openslr63.
