datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.Treble10-Speech
Treble10-Speech (16 kHz)
The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.sberdevices_golos_10h_crowd
Dataset Card for sberdevices_golos_10h_crowd
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.vlsp2020_vinai_100h
unofficial mirror of VLSP 2020 - VinAI - ASR challenge dataset
official announcement:
tiếng việt: https://institute.vinbigdata.org/events/vinbigdata-chia-se-100-gio-du-lieu-tieng-noi-cho-cong-dong/
in eglish: https://institute.vinbigdata.org/en/events/vinbigdata-shares-100-hour-data-for-the-community/
VLSP 2020 workshop: https://vlsp.org.vn/vlsp2020
official download: https://drive.google.com/file/d/1vUSxdORDxk-ePUt-bUVDahpoXiqKchMx/view?usp=sharing
contact: info@vinbigdata.org… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vlsp2020_vinai_100h.sberdevices_golos_100h_farfield
Dataset Card for sberdevices_golos_100h_farfield
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.bangla-10k
Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh
Bangla-10K is a 10,070.8-hour Bengali speech corpus with
567,323 recordings from India and Bangladesh. It combines scripted
single-speaker read speech with natural multi-speaker conversations for
Bengali automatic speech recognition (ASR). The paper rounds the corpus scale
to 10,000 hours.
The corpus and its ASR evaluation are described in the anonymous manuscript… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.Lingala_100hrs
Lingala 100hrs
110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated
from three publicly available CC-BY-4.0 corpora for ASR research.
Composition
Counts from a full-pass audit on 2026-07-09:
Source
Upstream location
Rows
Splits
AfriVoice (Lingala)
https://huggingface.co/datasets/DigitalUmuganda/AfriVoice
17,544
train (16,144), validation (915), test (485)
LRSC (Lingala Read Speech Corpus)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.libris_clean_100
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/libris_clean_100.kinyarwanda-speech-1000h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~1000 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.librispeech-alignments_clean100
librispeech-alignments_clean100
This is a subset of librispeech-alignments (https://huggingface.co/datasets/gilkeyio/librispeech-alignments) which only includes train_clean_100 and test_clean splits for small experiments and tutorials.
Cite:
@inproceedings{panayotov2015librispeech,
title={Librispeech: an ASR corpus based on public domain audio books},
author={Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev},
booktitle={ICASSP},
year={2015}… See the full description on the dataset page: https://huggingface.co/datasets/ErfanAShams/librispeech-alignments_clean100.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.1000h-us-english-smartphone-conversation
📚 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings)
This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for:
Automatic Speech Recognition (ASR)
Speaker Identification and Gender/Age Analysis
Dialect and Accent Modeling
Multi-speaker Speech Separation
🧾 Dataset Contents
The dataset includes:
metadata.CSV: Metadata including speaker gender, age… See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.radiotalk-us-transcripts-qwen3-100k
radiotalk-us-transcripts-qwen3-100k
100,000 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the
radiotalk transcripts series; the v2 release with higher per-transcript
realism lives at
twangodev/radiotalk-us-transcripts-qwen3-25k.
Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the
generator model in the dataset name; the old id redirects here.
Irodori-Ja-Spk4-10k
SynDataLab/Irodori-Ja-Spk4-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk4): 30s female, news-anchor mature — 30代女性、ニュースキャスター風の落ち着いた声.
How this speaker was made
The voice identity for Spk4 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk4-10k.open-vi-dialog-synthetic-100h
OpenDialog Vietnamese Synthetic Dialogue 100h
Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments.
12,000 chunks
30 seconds per chunk
100.0 hours total
Each item contains S1/S2 speaker labels, turn timings, target text,
relationship, pronouns, environment, topic, mood, and source reference IDs.
Audio renderer: vLLM-Omni VoxCPM2
Audio format: mono WAV, 48 kHz, 30 seconds per chunk
This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.Irodori-Ja-Spk3-10k
SynDataLab/Irodori-Ja-Spk3-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声.
How this speaker was made
The voice identity for Spk3 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k.Hypa-Speech-10k
A multilingual instruction-tuning dataset covering translation,transcription, and language detection.
Dataset Card for Hypa-Speech-10k
Dataset Summary
Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets.
The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Speech-10k.Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.Irodori-Ja-Spk1-10k
SynDataLab/Irodori-Ja-Spk1-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk1): 30s male, calm conversational — 30代男性、落ち着いた自然な会話調.
How this speaker was made
The voice identity for Spk1 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk1-10k.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/linq1005/fleurs.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/mls-eng-10k-tags_tagged_10k_generated.Hypa-Speech-10k
A multilingual instruction-tuning dataset covering translation,transcription, and language detection.
Dataset Card for Hypa-Speech-10k
Dataset Summary
Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets.
The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/Hypa-Speech-10k.vais1000
unofficial mirror of VAIS-1000
official announcement: https://vais.vn/vi/tai-ve/hts_for_vietnamese (dead)
mirror: https://github.com/undertheseanlp/text_to_speech/tree/run/data/vais1000/raw
small only 1h40min audio - 1 speaker (female northern accent) - 1k samples
pre-process: none
need to do: check misspelling, restore foreign words phonetised to vietnamese
usage with HuggingFace:
# pip install -q "datasets[audio]"
from datasets import load_dataset
from torch.utils.data import… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vais1000.Irodori-Ja-Spk2-10k
SynDataLab/Irodori-Ja-Spk2-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk2): 30s female, narrator-style natural — 30代女性、ナレーター風の自然な声.
How this speaker was made
The voice identity for Spk2 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk2-10k.cv10-uk-testset-clean
The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦
Overview
This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios.
All audios have been checked by a human to be sure that they are correct.
This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk
Community
Discord: https://bit.ly/discord-uds
Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.ViMD_chunked_10s
ViMD Chunked 10s — 16kHz
Preprocessed from ViMD (Nguyen et al., EMNLP 2024).
Preprocessing
Resample: 44.1kHz -> 16kHz mono
Chunking: each audio is split into consecutive NON-OVERLAPPING segments
of at most 10 seconds. ALL segments are kept, including the final
remainder (no minimum length filter). 1 original file -> ceil(len/10s) samples.
Splits: original ViMD train/valid/test kept (speaker-exclusive).
Segments of the same file always stay in the same split.… See the full description on the dataset page: https://huggingface.co/datasets/tannhoo06/ViMD_chunked_10s.taiwan-conversation-context-100-domains
Taiwan Conversation Context 100 Domains
Dataset Description
Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。
本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。
資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於:
語音生成資料前處理
Text-to-Speech, TTS
Spoken Dialogue Generation
Conversational AI
Customer Service Dialogue Modeling
Role-play Dialogue Dataset
台灣繁體中文語音模型訓練
生活情境問答模型訓練
對話式 AI 助理訓練
RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.
