datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.Treble10-Speech
Treble10-Speech (16 kHz)
The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.sberdevices_golos_10h_crowd
Dataset Card for sberdevices_golos_10h_crowd
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.vlsp2020_vinai_100h
unofficial mirror of VLSP 2020 - VinAI - ASR challenge dataset
official announcement:
tiếng việt: https://institute.vinbigdata.org/events/vinbigdata-chia-se-100-gio-du-lieu-tieng-noi-cho-cong-dong/
in eglish: https://institute.vinbigdata.org/en/events/vinbigdata-shares-100-hour-data-for-the-community/
VLSP 2020 workshop: https://vlsp.org.vn/vlsp2020
official download: https://drive.google.com/file/d/1vUSxdORDxk-ePUt-bUVDahpoXiqKchMx/view?usp=sharing
contact: info@vinbigdata.org… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vlsp2020_vinai_100h.bangla-10k
Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh
Bangla-10K is a 10,816-hour Bengali speech corpus with
624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus
(567,323 recordings) and a separately collected 745.1-hour evaluation set
(57,628 recordings). It combines scripted single-speaker read speech with
natural multi-speaker conversations for Bengali automatic speech recognition
(ASR).
The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.sberdevices_golos_100h_farfield
Dataset Card for sberdevices_golos_100h_farfield
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.Lingala_100hrs
Lingala 100hrs
110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated
from three publicly available CC-BY-4.0 corpora for ASR research.
Composition
Counts from a full-pass audit on 2026-07-09:
Source
Upstream location
Rows
Splits
AfriVoice (Lingala)
https://huggingface.co/datasets/DigitalUmuganda/AfriVoice
17,544
train (16,144), validation (915), test (485)
LRSC (Lingala Read Speech Corpus)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.libris_clean_100
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/libris_clean_100.malagasy-asr-100h-split
Malagasy ASR 100h Split
This dataset contains the cleaned and reviewed Malagasy ASR 100h split.
Splits
Split
Samples
Hours
train
33505
90.1227
validation
1860
4.9134
test
1860
4.9640
Columns
audio: relative path to the audio file
text: reviewed reference transcription
duration_sec: audio duration in seconds
source: source label (waxal or voxlingua)
Notes
This split includes the final human review corrections… See the full description on the dataset page: https://huggingface.co/datasets/malagasy-asr/malagasy-asr-100h-split.kinyarwanda-speech-1000h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~1000 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.librispeech-alignments_clean100
librispeech-alignments_clean100
This is a subset of librispeech-alignments (https://huggingface.co/datasets/gilkeyio/librispeech-alignments) which only includes train_clean_100 and test_clean splits for small experiments and tutorials.
Cite:
@inproceedings{panayotov2015librispeech,
title={Librispeech: an ASR corpus based on public domain audio books},
author={Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev},
booktitle={ICASSP},
year={2015}… See the full description on the dataset page: https://huggingface.co/datasets/ErfanAShams/librispeech-alignments_clean100.huper-clean100-proxyphones
huper-clean100-proxyphones
LibriSpeech train-clean-100 audio paired with HuPER-style proxy ARPAbet phone labels (machine-generated / proxy, not human verified).This corresponds to the 100h train-clean-100 split (28,539 utterances).Note: HuggingFace Dataset Viewer is not supported because the data is provided as tar+zstd shards. Follow the instructions below to download and extract locally.
What’s inside
The data is stored as 5 shards under blobs/:… See the full description on the dataset page: https://huggingface.co/datasets/huper29/huper-clean100-proxyphones.1000h-us-english-smartphone-conversation
📚 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings)
This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for:
Automatic Speech Recognition (ASR)
Speaker Identification and Gender/Age Analysis
Dialect and Accent Modeling
Multi-speaker Speech Separation
🧾 Dataset Contents
The dataset includes:
metadata.CSV: Metadata including speaker gender, age… See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.Irodori-Ja-Spk4-10k
SynDataLab/Irodori-Ja-Spk4-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk4): 30s female, news-anchor mature — 30代女性、ニュースキャスター風の落ち着いた声.
How this speaker was made
The voice identity for Spk4 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk4-10k.Irodori-Ja-Spk3-10k
SynDataLab/Irodori-Ja-Spk3-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声.
How this speaker was made
The voice identity for Spk3 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k.open-vi-dialog-synthetic-100h
OpenDialog Vietnamese Synthetic Dialogue 100h
Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments.
12,000 chunks
30 seconds per chunk
100.0 hours total
Each item contains S1/S2 speaker labels, turn timings, target text,
relationship, pronouns, environment, topic, mood, and source reference IDs.
Audio renderer: vLLM-Omni VoxCPM2
Audio format: mono WAV, 48 kHz, 30 seconds per chunk
This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.Hypa-Speech-10k
A multilingual instruction-tuning dataset covering translation,transcription, and language detection.
Dataset Card for Hypa-Speech-10k
Dataset Summary
Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets.
The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Speech-10k.Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.Irodori-Ja-Spk1-10k
SynDataLab/Irodori-Ja-Spk1-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk1): 30s male, calm conversational — 30代男性、落ち着いた自然な会話調.
How this speaker was made
The voice identity for Spk1 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk1-10k.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/linq1005/fleurs.Hypa-Speech-10k
A multilingual instruction-tuning dataset covering translation,transcription, and language detection.
Dataset Card for Hypa-Speech-10k
Dataset Summary
Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets.
The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/Hypa-Speech-10k.vais1000
unofficial mirror of VAIS-1000
official announcement: https://vais.vn/vi/tai-ve/hts_for_vietnamese (dead)
mirror: https://github.com/undertheseanlp/text_to_speech/tree/run/data/vais1000/raw
small only 1h40min audio - 1 speaker (female northern accent) - 1k samples
pre-process: none
need to do: check misspelling, restore foreign words phonetised to vietnamese
usage with HuggingFace:
# pip install -q "datasets[audio]"
from datasets import load_dataset
from torch.utils.data import… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vais1000.Irodori-Ja-Spk2-10k
SynDataLab/Irodori-Ja-Spk2-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk2): 30s female, narrator-style natural — 30代女性、ナレーター風の自然な声.
How this speaker was made
The voice identity for Spk2 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk2-10k.cv10-uk-testset-clean
The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦
Overview
This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios.
All audios have been checked by a human to be sure that they are correct.
This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk
Community
Discord: https://bit.ly/discord-uds
Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.ViMD_chunked_10s
ViMD Chunked 10s — 16kHz
Preprocessed from ViMD (Nguyen et al., EMNLP 2024).
Preprocessing
Resample: 44.1kHz -> 16kHz mono
Chunking: each audio is split into consecutive NON-OVERLAPPING segments
of at most 10 seconds. ALL segments are kept, including the final
remainder (no minimum length filter). 1 original file -> ceil(len/10s) samples.
Splits: original ViMD train/valid/test kept (speaker-exclusive).
Segments of the same file always stay in the same split.… See the full description on the dataset page: https://huggingface.co/datasets/tannhoo06/ViMD_chunked_10s.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.luganda_bible_audio_100Hrs
Luganda Bible Audio 100Hrs
Audio split by chapters of Old & New Testament & Transcripts
Format: 64kps, MP3
Requires further splitting of the audio into smaller chunks/splits (https://github.com/facebookresearch/fairseq/tree/main/examples/mms/data_prep)
Audio sourced from https://www.faithcomesbyhearing.com/
minds14
MInDS-14
MINDS-14 is training and evaluation resource for intent detection task with spoken data. It covers 14
intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties.
Example
MInDS-14 can be downloaded and used as follows:
from datasets import load_dataset
minds_14 = load_dataset("PolyAI/minds14", "fr-FR") # for French
# to download all data for multi-lingual fine-tuning uncomment… See the full description on the dataset page: https://huggingface.co/datasets/yxl10086/minds14.
