datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LEMAS-Dataset-train
Overview
This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.Thinkspark-v2-270m-training-data
ThinkSpark-v2-350M — training data
Full-duplex floor-controller (Section 8) training corpus: playable audio + text,
paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain,
gender, prosody, agent text) and Soniox character-level timestamps.
Dataset Viewer
Default split is parquet with a real Audio feature — a player renders inline next to
the text in the Hub UI:
column
type
description
audio
Audio
playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.knesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.ellipsis-lrs3-trainval-autoavsr
LRS3 trainval — auto_avsr mouth crops, bridged to WebDataset
The LRS3 trainval split (30 h) as [T,96,96] uint8 grayscale mouth-ROI shards with character
transcripts, produced with auto_avsr's own crop geometry so results stay comparable to that
project's published LRS3 numbers.
Derived from TheNHz/ellipsis-lrs3-raw
(itself a verified mirror of LRS3-TED, whose official distribution was discontinued).
Attribution
LRS3-TED is by Triantafyllos Afouras, Joon Son Chung… See the full description on the dataset page: https://huggingface.co/datasets/TheNHz/ellipsis-lrs3-trainval-autoavsr.crowd-recital-whisper-training
Dataset Card for ivrit.ai - Crowd Recital
Dataset Details
Dataset Description
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
- Full license: https://www.ivrit.ai/en/the-license/
- FAQs: https://www.ivrit.ai/en/license-faqs/
Dataset Structure
Data Fields
Each example in the dataset contains:
audio: An audio column containing:
bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.lao_stt_training_data
Lao Speech-to-Text Training Data
ຊຸດຂໍ້ມູນນີ້ຖືກຈັດກຽມຂຶ້ນມາເພື່ອໃຊ້ສຳລັບການເທຣນ ແລະ ປັບແຕ່ງ (Fine-tuning) ໂມເດວ Speech-to-Text (ເຊັ່ນ OpenAI Whisper) ສຳລັບພາສາລາວ.
ໂຄງສ້າງຂອງຂໍ້ມູນ (Dataset Structure)
Train set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ train/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ train.csv
Validation set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ validation/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ validation.csv
ຮູບແບບຂໍ້ມູນໃນໄຟລ໌ CSV:
audio: ເສັ້ນທາງໄປຫາໄຟລ໌ສຽງ (e.g., train/audio25000.wav)… See the full description on the dataset page: https://huggingface.co/datasets/KitTzk/lao_stt_training_data.pseudolabel-malaya-speech-stt-train-whisper-large-v3ai-auto-train-datasets-cuda-5d-quantum-mindmap-simulations-generator-zkevms-immutablexuzbek-asr-train-manifests
Uzbek ASR Training Manifests
The exact training, validation and test splits behind
rustam1221/uzbek-asr-gigaam:
974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized,
and split by speaker.
No audio is copied. Each row is a pointer — a parquet file plus a row
index in the upstream dataset — and the training dataloader decodes the audio
when the batch is built. That keeps the whole corpus definition at 200 MB
instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.libristutter_trainYou may use this dataset to finetune a WhisperingLLaMa style model on stuttering data from the LibriStutter dataset. We have cleaned the data to work specifically for this task with hat model type.
new_train_data
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/developeranalyser/new_train_data.tts-training-dataset
Human Reviewed Telugu-English TTS Dataset
A manually reviewed multilingual TTS dataset created from publicly available educational and speech content.
Dataset Splits & Distribution Metrics
balanced_60min Split
Total Segments: 120
Total Duration: 60.00 minutes
Unique Speakers: 3
Distribution Breakdowns:
Language Distribution:
en-IN: 60 segments (30.00 minutes)
te-IN: 60 segments (30.00 minutes)
Style Distribution:
analytical: 27… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tts-training-dataset.whisper-large-training
Khmer Speech Dataset for Whisper Large V3
Combined Khmer speech datasets for fine-tuning Whisper models.
Statistics
Total: 52,255 samples (~47.3 hours)
Train: 49,119 (94.0%)
Test: 3,136 (6.0%)
Duration: 1.0s - 13.3s (avg: 3.3s)
Sources
seanghay/khmer_mpwt_speech (×5)
Samples: 10,290 (duplicated 5x from 2,058)
Duration: ~9.6 hours
seanghay/km-speech-corpus
Samples: 14,943
Duration: ~10.3 hours
google/fleurs
Samples: 1,675… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/whisper-large-training.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/0x3/Easy-Turn-Trainset.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/KikuyuASR_trainingdataset.isizulu-asr-train
isiZulu Speech Recognition Augmented Train Dataset
Dataset Description
This dataset contains augmented speech recordings and transcriptions for isiZulu, one of South Africa's official languages.
The dataset has been optimized for use with OpenAI's Whisper ASR models.
Dataset Statistics
Number of samples: 690
Language: isiZulu (Zul)
Audio format: WAV, 16kHz, mono, 16-bit
Maximum duration: 30 seconds (truncated for Whisper compatibility)
Transcription format:… See the full description on the dataset page: https://huggingface.co/datasets/zionia/isizulu-asr-train.trainSetcrowd-recital-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Recital - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~78h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.crowd-whatsapp-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~19h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.quantized-librispeech-train-360isixhosa-asr-train
isiXhosa Speech Recognition Augmented Dataset
Dataset Description
This dataset contains augmented speech recordings and transcriptions for isiXhosa, one of South Africa's official languages.
The dataset has been optimized for use with OpenAI's Whisper ASR models.
Dataset Statistics
Number of samples: 680
Language: isiXhosa (Xho)
Audio format: WAV, 16kHz, mono, 16-bit
Maximum duration: 30 seconds (truncated for Whisper compatibility)
Transcription format:… See the full description on the dataset page: https://huggingface.co/datasets/zionia/isixhosa-asr-train.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/jo-05/KikuyuASR_trainingdataset.kazakh_speech_corpus_2
Kazakh_speech_dataset_2
This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format.
Dataset info
645,860 Utterances
1194 Hours in total
Sources in each split:
test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'}
train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'}
validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'}
Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.kazakh_speech_dataset_ksdKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api.
Dataset info:
813 Speakers
with 500 samples for 4 speakers
with 250 samples for 809 speakers
Male/female
555 Hours
Guides
Load data 1
Replace the export HF_HOME with your HF_HOME path
from datasets import load_dataset
# export HF_HOME="/data/vladimir_albrekht/hf_cache"
ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.TrainingSpeechTrainingSpeech is an initiative to provide open and freely reusable dataset of voices
for speech-to-text models training
on non-english languages
using already available data (such as audio-books).
Right now, data are extracted exclusively from audio-books and in French language. Let me know if you are intersted to contribute by creating an issue.
Tooling
TrainingSpeech comes with a CLI that automate and simplify:
transcript extraction
forced-alignment (using aeneas)… See the full description on the dataset page: https://huggingface.co/datasets/wasertech/TrainingSpeech.AdoCleanCode_french-multi-mfa_train_v1_previewCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données AdoCleanCode/french-multi-mfa_train_v1_preview .
KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/KikuyuASR_trainingdataset.
