datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Thinkspark-v2-270m-training-data
ThinkSpark-v2-350M — training data
Full-duplex floor-controller (Section 8) training corpus: playable audio + text,
paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain,
gender, prosody, agent text) and Soniox character-level timestamps.
Dataset Viewer
Default split is parquet with a real Audio feature — a player renders inline next to
the text in the Hub UI:
column
type
description
audio
Audio
playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.lao_stt_training_data
Lao Speech-to-Text Training Data
ຊຸດຂໍ້ມູນນີ້ຖືກຈັດກຽມຂຶ້ນມາເພື່ອໃຊ້ສຳລັບການເທຣນ ແລະ ປັບແຕ່ງ (Fine-tuning) ໂມເດວ Speech-to-Text (ເຊັ່ນ OpenAI Whisper) ສຳລັບພາສາລາວ.
ໂຄງສ້າງຂອງຂໍ້ມູນ (Dataset Structure)
Train set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ train/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ train.csv
Validation set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ validation/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ validation.csv
ຮູບແບບຂໍ້ມູນໃນໄຟລ໌ CSV:
audio: ເສັ້ນທາງໄປຫາໄຟລ໌ສຽງ (e.g., train/audio25000.wav)… See the full description on the dataset page: https://huggingface.co/datasets/KitTzk/lao_stt_training_data.tts-training-dataset
Human Reviewed Telugu-English TTS Dataset
A manually reviewed multilingual TTS dataset created from publicly available educational and speech content.
Dataset Splits & Distribution Metrics
balanced_60min Split
Total Segments: 120
Total Duration: 60.00 minutes
Unique Speakers: 3
Distribution Breakdowns:
Language Distribution:
en-IN: 60 segments (30.00 minutes)
te-IN: 60 segments (30.00 minutes)
Style Distribution:
analytical: 27… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tts-training-dataset.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/KikuyuASR_trainingdataset.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/jo-05/KikuyuASR_trainingdataset.kazakh_speech_dataset_ksdKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api.
Dataset info:
813 Speakers
with 500 samples for 4 speakers
with 250 samples for 809 speakers
Male/female
555 Hours
Guides
Load data 1
Replace the export HF_HOME with your HF_HOME path
from datasets import load_dataset
# export HF_HOME="/data/vladimir_albrekht/hf_cache"
ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/KikuyuASR_trainingdataset.
