datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Thinkspark-v2-270m-training-data
ThinkSpark-v2-350M — training data
Full-duplex floor-controller (Section 8) training corpus: playable audio + text,
paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain,
gender, prosody, agent text) and Soniox character-level timestamps.
Dataset Viewer
Default split is parquet with a real Audio feature — a player renders inline next to
the text in the Hub UI:
column
type
description
audio
Audio
playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.knesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.crowd-recital-whisper-training
Dataset Card for ivrit.ai - Crowd Recital
Dataset Details
Dataset Description
License
The dataset is released under the ivrit.ai License, which enables broad research and commercial use.
- Full license: https://www.ivrit.ai/en/the-license/
- FAQs: https://www.ivrit.ai/en/license-faqs/
Dataset Structure
Data Fields
Each example in the dataset contains:
audio: An audio column containing:
bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.lao_stt_training_data
Lao Speech-to-Text Training Data
ຊຸດຂໍ້ມູນນີ້ຖືກຈັດກຽມຂຶ້ນມາເພື່ອໃຊ້ສຳລັບການເທຣນ ແລະ ປັບແຕ່ງ (Fine-tuning) ໂມເດວ Speech-to-Text (ເຊັ່ນ OpenAI Whisper) ສຳລັບພາສາລາວ.
ໂຄງສ້າງຂອງຂໍ້ມູນ (Dataset Structure)
Train set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ train/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ train.csv
Validation set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ validation/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ validation.csv
ຮູບແບບຂໍ້ມູນໃນໄຟລ໌ CSV:
audio: ເສັ້ນທາງໄປຫາໄຟລ໌ສຽງ (e.g., train/audio25000.wav)… See the full description on the dataset page: https://huggingface.co/datasets/KitTzk/lao_stt_training_data.tts-training-dataset
Human Reviewed Telugu-English TTS Dataset
A manually reviewed multilingual TTS dataset created from publicly available educational and speech content.
Dataset Splits & Distribution Metrics
balanced_60min Split
Total Segments: 120
Total Duration: 60.00 minutes
Unique Speakers: 3
Distribution Breakdowns:
Language Distribution:
en-IN: 60 segments (30.00 minutes)
te-IN: 60 segments (30.00 minutes)
Style Distribution:
analytical: 27… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tts-training-dataset.whisper-large-training
Khmer Speech Dataset for Whisper Large V3
Combined Khmer speech datasets for fine-tuning Whisper models.
Statistics
Total: 52,255 samples (~47.3 hours)
Train: 49,119 (94.0%)
Test: 3,136 (6.0%)
Duration: 1.0s - 13.3s (avg: 3.3s)
Sources
seanghay/khmer_mpwt_speech (×5)
Samples: 10,290 (duplicated 5x from 2,058)
Duration: ~9.6 hours
seanghay/km-speech-corpus
Samples: 14,943
Duration: ~10.3 hours
google/fleurs
Samples: 1,675… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/whisper-large-training.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/KikuyuASR_trainingdataset.crowd-recital-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Recital - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~78h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.crowd-whatsapp-yi-whisper-training
Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish
See more details on the source dataset card.
Dataset Details
Dataset Description
This is a derived dataset for structured for whisper training:
Excludes low quality segments (judged by probabilities of the text-audio auto alignment process)
Encodes timestamps along segments of text + previous text
Audio encoded to 16K sample-rate, mono
Total audio duration - ~19h
License: other
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/jo-05/KikuyuASR_trainingdataset.kazakh_speech_corpus_2
Kazakh_speech_dataset_2
This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format.
Dataset info
645,860 Utterances
1194 Hours in total
Sources in each split:
test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'}
train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'}
validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'}
Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.kazakh_speech_dataset_ksdKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api.
Dataset info:
813 Speakers
with 500 samples for 4 speakers
with 250 samples for 809 speakers
Male/female
555 Hours
Guides
Load data 1
Replace the export HF_HOME with your HF_HOME path
from datasets import load_dataset
# export HF_HOME="/data/vladimir_albrekht/hf_cache"
ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.TrainingSpeechTrainingSpeech is an initiative to provide open and freely reusable dataset of voices
for speech-to-text models training
on non-english languages
using already available data (such as audio-books).
Right now, data are extracted exclusively from audio-books and in French language. Let me know if you are intersted to contribute by creating an issue.
Tooling
TrainingSpeech comes with a CLI that automate and simplify:
transcript extraction
forced-alignment (using aeneas)… See the full description on the dataset page: https://huggingface.co/datasets/wasertech/TrainingSpeech.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/KikuyuASR_trainingdataset.
